Methodology
Our tools answer one question: what’s the right model for your task at the lowest cost? To do that honestly, we’re explicit about where the numbers come from and what they do and don’t mean.
Where the intelligence signal comes from
Quality scores are community Arena ratings from LMArena / arena.ai: people compare two anonymous model answers side by side and pick the better one, producing an Elo-style score. A higher score means a model is preferred more often in blind head-to-head votes. It reflects human preference, not correctness on any single benchmark, and it can favour style as well as substance.
We currently carry scores for 114 models across these arenas: Overall, Coding, Math & reasoning, Agentic / tool use. Prices come from each provider’s published rates and are refreshed daily (see any model page for its source and last-changed date).
The full ranking is published as open data: intelligence.json and intelligence.csv. Reuse it with attribution and a link back — see the leaderboard for the citation.
Intelligence per dollar
On the leaderboard we divide a model’s Arena score by its blended price — a 3:1 input:output blend in dollars per 1M tokens, the same blend used across the site. It surfaces models that punch above their price. It is a starting point, not a verdict: a tiny model with a great ratio may still be below the quality bar your task needs.
The quality bar
The smartest model is rarely the one you should ship. Our recommender takes every model whose Arena score is within a chosen percentage of the best in that arena, filters by the capabilities and context you require, and returns the cheapest survivor. You choose how close to the top you need to be — “good enough” is a legitimate, often much cheaper, answer.
The price/quality frontier
A model is on the frontier if nothing else is both at least as good and at least as cheap. Models we mark Best value are on this frontier — everything off it is beaten by something on it. This is the honest way to compare across price tiers without inventing a single opaque “value score”.
Splitting work across models
Most workloads mix a few hard calls with a lot of routine ones. The plan smart, build cheap calculator estimates the savings from routing hard calls to a strong model and the bulk to a cheap one, versus paying flagship prices for everything.
Measured vs estimated scores
We label every score. Measured means the Arena leaderboard voted on that exact model. Estimated means the model is a regional or creator-prefixed hosting listing of a base model — for example an AWS Bedrock jp.anthropic.… inference profile. The weights are identical, so we inherit the base model’s score, but Arena never voted on that specific listing. We surface the label rather than hide the model, so you always know whether a score is direct or inherited.
Why we fail toward silence
Matching a leaderboard’s model names to the exact API models we price is error-prone: variants, dated snapshots and reasoning-effort tags all muddy the join. Our rule is fail-closed — if we can’t confidently map a score to a model, we show no score rather than a wrong one. That’s why some models have no intelligence data, and why search-specialised models (e.g. Perplexity Sonar) are excluded from the general arenas.
Limitations
- Arena scores measure preference, not task-specific accuracy — always validate on your own data.
- Scores drift as models are added and votes accumulate; we snapshot them and note the date.
- Cost estimates assume the token mix you enter and standard (non-discounted) pricing.
- We are independent and take no payment from providers to influence rankings.
Intelligence scores are community Arena ratings from LMArena / arena.ai, used under CC BY 4.0. Snapshot last refreshed 28 July 2026. Scores are a relative signal, not an absolute measure of capability. A Measured score was voted on directly; an Estimated score is inherited from an identical base model (a regional/creator-prefixed hosting duplicate). See our methodology.