AI Calculator Pro

Methodology

Our tools answer one question: what’s the right model for your task at the lowest cost? To do that honestly, we’re explicit about where the numbers come from and what they do and don’t mean.

What we track, and what we list

A model directory can be big or it can be trustworthy. We see prices for 1,768 models across 6,750 provider offerings, and we list 347 of them. The rest are deliberately left out: deprecated base models, dated snapshots, preview builds and fine-tune variants inflate a headline number and tell you nothing about what to actually buy.

Models sold directly by a lab we track — OpenAI, Anthropic, Google, Mistral and the rest — are listed on that basis alone. The filters below decide the harder case: open-weight models that no lab sells, which exist only as weights resold by inference hosts. That is where a directory either stays honest or fills up with noise. Recomputed on every daily rebuild:

ModelsStageWhy
1,768Models we see a price forEvery distinct model priced by at least one provider.
420Sold by 3+ independent providersOne reseller listing a model proves nothing. Several providers carrying the same weights is a demand signal we can't fake.
355Made by a lab we trackWe need to know who built a model to describe it honestly.
252Passed our quality filtersWe drop preview builds, dated snapshots, base (non-instruct) models, fine-tune variants and anything priced outside sane bounds.

Those survivors plus the first-party models are what make up the 347 we list. Every one is searchable, comparable and included in our open datasets. A smaller set — 71 — has been through human review and carries a page we submit to search engines; the rest are labelled auto-discovered and stay out of the index until someone has checked them.

We would rather be the shorter, defensible list. Every number above is emitted by the pipeline that applies the filters, not typed into a marketing line someone has to remember to update — so if the bar moves, the page moves with it.

Where the intelligence signal comes from

Quality scores are community Arena ratings from LMArena / arena.ai: people compare two anonymous model answers side by side and pick the better one, producing an Elo-style score. A higher score means a model is preferred more often in blind head-to-head votes. It reflects human preference, not correctness on any single benchmark, and it can favour style as well as substance.

We currently carry scores for 158 models across these arenas: Overall, Coding, Math & reasoning, Agentic / tool use. Prices come from each provider’s published rates and are refreshed daily (see any model page for its source and last-changed date).

The full ranking is published as open data: intelligence.json and intelligence.csv. Reuse it with attribution and a link back — see the leaderboard for the citation.

Intelligence per dollar

On the leaderboard we divide a model’s Arena score by its blended price — a 3:1 input:output blend in dollars per 1M tokens, the same blend used across the site. It surfaces models that punch above their price. It is a starting point, not a verdict: a tiny model with a great ratio may still be below the quality bar your task needs.

The quality bar

The smartest model is rarely the one you should ship. Our recommender takes every model whose Arena score is within a chosen percentage of the best in that arena, filters by the capabilities and context you require, and returns the cheapest survivor. You choose how close to the top you need to be — “good enough” is a legitimate, often much cheaper, answer.

The price/quality frontier

A model is on the frontier if nothing else is both at least as good and at least as cheap. Models we mark Best value are on this frontier — everything off it is beaten by something on it. This is the honest way to compare across price tiers without inventing a single opaque “value score”.

Splitting work across models

Most workloads mix a few hard calls with a lot of routine ones. The plan smart, build cheap calculator estimates the savings from routing hard calls to a strong model and the bulk to a cheap one, versus paying flagship prices for everything.

Measured vs estimated scores

We label every score. Measured means the Arena leaderboard voted on that exact model. Estimated means the model is a regional or creator-prefixed hosting listing of a base model — for example an AWS Bedrock jp.anthropic.… inference profile. The weights are identical, so we inherit the base model’s score, but Arena never voted on that specific listing. We surface the label rather than hide the model, so you always know whether a score is direct or inherited.

Why we fail toward silence

Matching a leaderboard’s model names to the exact API models we price is error-prone: variants, dated snapshots and reasoning-effort tags all muddy the join. Our rule is fail-closed — if we can’t confidently map a score to a model, we show no score rather than a wrong one. That’s why some models have no intelligence data, and why search-specialised models (e.g. Perplexity Sonar) are excluded from the general arenas.

Limitations

  • Arena scores measure preference, not task-specific accuracy — always validate on your own data.
  • Scores drift as models are added and votes accumulate; we snapshot them and note the date.
  • Cost estimates assume the token mix you enter and standard (non-discounted) pricing.
  • We are independent and take no payment from providers to influence rankings.

Intelligence scores are community Arena ratings from LMArena / arena.ai, used under CC BY 4.0. Snapshot last refreshed 18 September 2026. Scores are a relative signal, not an absolute measure of capability. A Measured score was voted on directly; an Estimated score is inherited from an identical base model (a regional/creator-prefixed hosting duplicate). See our methodology.