LLM Pricing Comparison
Compare input, output and cached prices across models.
With the default inputs, LLM pricing compared — 326 models, per 1M tokens. Enter your own numbers below to recompute instantly; the full step-by-step math is shown under the worked example.
Compare the input, output and cached-input prices of every model in one ranked table, cheapest first, to find the best value for your workload.
How this is calculated
We rank every model by input, output and cached-input price per 1M tokens, cheapest first, so you can see raw price positioning at a glance. Unlike the cost comparison, this ignores your specific workload and just compares list prices.
Is this a good result? What to do next
Lowest list price isn't lowest cost for your workload — output-heavy tasks weight the output price, input-heavy tasks the input price. Use this for a quick market read, then the cost comparison for your actual token mix.
Typical planning ranges
- Cheapest
- Nova Micro, Command R7B, Gemini Flash-Lite, DeepSeek V4 Flash
- Output
- usually 3–5× input
- Cached input
- a fraction of normal input
Ranges are typical planning figures to sanity-check your result, not authoritative benchmarks. Your numbers will vary with use case, volume, and vendor.
How to improve this number
- Weigh input vs output price by your workload's mix.
- Watch cached-input pricing for cache-heavy workloads.
- Shortlist on price, then confirm quality.
Common mistakes
- Ranking on input price alone when output dominates.
- Ignoring cached-input rates for RAG/system-prompt-heavy apps.
When to use a different approach
For cost on your real workload, use the model cost comparison. For a head-to-head, use a two-model comparison.
Worked example (defaults)
With the default inputs above, here is the result:
Sources & references
Frequently asked questions
Which LLM is cheapest per token?+
Open-weight and small models (Nova Micro, Command R7B, Gemini Flash-Lite, DeepSeek V4 Flash) are cheapest. The table ranks all models by combined input+output price.