AI Calculator Pro

Cheapest model that fits a 200k-token context

Teams with very long prompts often need to find the cheapest model that fits a 200k-token context. Because providers price input and output tokens very differently, the "obvious" model is rarely the cheapest. This tool compares them on your own workload so the decision is data-driven.

Start here: Model Context Window Comparison

Result
Context windows compared
223 models ranked by context size
Llama 4 ScoutMeta10,000,0008,192
GPT-5.4OpenAI1,050,000128,000
GPT-5.4 ProOpenAI1,050,000128,000
GPT-5.5OpenAI1,050,000128,000
GPT-5.5 ProOpenAI1,050,000128,000
GPT-5.6OpenAI1,050,000128,000
GPT-5.6 LunaOpenAI1,050,000128,000
GPT-5.6 SolOpenAI1,050,000128,000
GPT-5.6 TerraOpenAI1,050,000128,000
Gemini 2.5 ProGoogle1,048,57665,536
Kimi K3Moonshot AI1,048,576131,072
MiMo-V2.5-ProXiaomi1,048,576131,072
MiMo-V2.5Xiaomi1,048,576131,072
Gemini 2.0 FlashGoogle1,048,5768,192
Gemini 2.0 Flash-LiteGoogle1,048,5768,192
Gemini 3.1 Flash LiteGoogle1,048,57665,536
Gemini 3.5 FlashGoogle1,048,57665,536
MiMo-V2.5-Pro-UltraSpeedXiaomi1,048,576131,072
MiMo-V2-ProXiaomi1,048,576131,072
Qwen3 Coder PlusAlibaba1,048,57665,536
Gemini 3.5 Flash LiteGoogle1,048,57665,536
Gemini 3.6 FlashGoogle1,048,57665,536
GPT-4.1OpenAI1,000,00032,768
GPT-4.1 MiniOpenAI1,000,00032,768
GPT-4.1 NanoOpenAI1,000,00032,768
Claude Opus 4.8Anthropic1,000,00064,000
Claude Opus 4.6Anthropic1,000,00064,000
Claude Sonnet 4.6Anthropic1,000,00064,000
Claude Sonnet 5Anthropic1,000,00064,000
Claude Fable 5Anthropic1,000,00064,000
Gemini 2.5 FlashGoogle1,000,00065,536
Gemini 2.5 Flash-LiteGoogle1,000,00065,536
Amazon Nova PremierAWS Bedrock1,000,00010,000
Llama 4 MaverickMeta1,000,0008,192
DeepSeek V4 FlashDeepSeek1,000,0008,192
DeepSeek V4 ProDeepSeek1,000,0008,192
Grok 4.3xAI1,000,000
Qwen3.6 PlusAlibaba1,000,00065,536
Qwen3.6 FlashAlibaba1,000,00065,536
GLM-5.2Z.ai (Zhipu AI)1,000,000131,072
MiniMax-M3MiniMax1,000,000128,000
Nemotron 3 Ultra 550B A55BNvidia1,000,000128,000
Claude Fable 5AWS Bedrock1,000,000128,000
Claude Opus 4.6 (US)AWS Bedrock1,000,000128,000
Claude Opus 4.7 (US)AWS Bedrock1,000,000128,000
Claude Opus 4.8AWS Bedrock1,000,000128,000
AU Anthropic Claude Sonnet 4.6AWS Bedrock1,000,000128,000
Claude Sonnet 5 (Global)AWS Bedrock1,000,000128,000
Claude Opus 4.7Anthropic1,000,000128,000
Claude Sonnet 4.5 (latest)Anthropic1,000,00064,000
DeepSeek ChatDeepSeek1,000,000384,000
DeepSeek ReasonerDeepSeek1,000,000384,000
Muse Spark 1.1Meta1,000,00032,000
Qwen FlashAlibaba1,000,00032,768
Qwen PlusAlibaba1,000,00032,768
Qwen TurboAlibaba1,000,00016,384
Qwen3.5 PlusAlibaba1,000,00065,536
Qwen3.7 MaxAlibaba1,000,00065,536
Qwen3.7 PlusAlibaba1,000,00065,536
Qwen3 Coder FlashAlibaba1,000,00065,536
Grok 4.3AWS Bedrock1,000,000131,072
Claude Opus 5 (JP)AWS Bedrock1,000,000128,000
Claude Opus 5Anthropic1,000,000128,000
Grok 4.5xAI500,000
GPT-5.1OpenAI400,000128,000
GPT-5.1 CodexOpenAI400,000128,000
GPT-5.1 Codex MaxOpenAI400,000128,000
GPT-5.1 Codex miniOpenAI400,000128,000
GPT-5.2OpenAI400,000128,000
GPT-5.2 CodexOpenAI400,000128,000
GPT-5.2 ProOpenAI400,000128,000
GPT-5.3 CodexOpenAI400,000128,000
GPT-5.4 miniOpenAI400,000128,000
GPT-5.4 nanoOpenAI400,000128,000
GPT-5-CodexOpenAI400,000128,000
GPT-5 NanoOpenAI400,000128,000
GPT-5 ProOpenAI400,000272,000
Amazon Nova ProAWS Bedrock300,0005,120
Amazon Nova LiteAWS Bedrock300,0005,120
GPT-5.4AWS Bedrock272,000128,000
GPT-5.5AWS Bedrock272,000128,000
GPT-5.6 LunaAWS Bedrock272,000128,000
GPT-5.6 SolAWS Bedrock272,000128,000
GPT-5.6 TerraAWS Bedrock272,000128,000
Kimi K2.7 CodeMoonshot AI262,144262,144
Qwen3 MaxAlibaba262,14465,536
Nemotron 3 Super 120B A12BNvidia262,144262,144
Nemotron 3 Nano 30B A3BNvidia262,14465,536
Hunyuan Hy3Tencent262,14464,000
Kimi K2.5Moonshot AI262,144262,144
Kimi K2.6Moonshot AI262,144262,144
Kimi K2.7 Code HighSpeedMoonshot AI262,144262,144
Kimi K2 ThinkingMoonshot AI262,144262,144
Kimi K2 Thinking TurboMoonshot AI262,144262,144
MiMo-V2-FlashXiaomi262,14465,536
MiMo-V2-OmniXiaomi262,144131,072
NVIDIA Nemotron 3 Super 120B A12BAWS Bedrock262,144131,072
Qwen3.5 122B-A10BAlibaba262,14465,536
Qwen3.5 27BAlibaba262,14465,536
Qwen3.5 35B-A3BAlibaba262,14465,536
Qwen3.5 397B-A17BAlibaba262,14465,536
Qwen3.6 27BAlibaba262,14465,536
Qwen3.6 35B-A3BAlibaba262,14465,536
Qwen3-Coder 30B-A3B InstructAlibaba262,14465,536
Qwen3-Coder 480B-A35B InstructAlibaba262,14465,536
Qwen3-VL PlusAlibaba262,14432,768
Kimi K2 ThinkingAWS Bedrock262,14316,000
Kimi K2.5AWS Bedrock262,14316,000
Mistral Large 3Mistral AI262,0008,192
Devstral 2Mistral AI262,0008,192
Qwen/Qwen3-Next-80B-A3B-InstructAWS Bedrock262,000262,000
Qwen/Qwen3-VL-235B-A22B-InstructAWS Bedrock262,000262,000
Command ACohere256,0008,000
MAI-Code-1-FlashMicrosoft256,000128,000
Step 3.7 FlashStepFun256,000256,000
Devstral 2 123BAWS Bedrock256,0008,192
Ministral 3 3BAWS Bedrock256,0008,192
Mistral Large 3AWS Bedrock256,0008,192
Step 3.5 FlashStepFun256,000256,000
GLM-5Z.ai (Zhipu AI)204,800131,072
MiniMax-M2.7MiniMax204,800131,072
GLM-4.6Z.ai (Zhipu AI)204,800131,072
GLM-4.7Z.ai (Zhipu AI)204,800131,072
MiniMax-M2.1MiniMax204,800131,072
MiniMax-M2.5MiniMax204,800131,072
MiniMax-M2.5-highspeedMiniMax204,800131,072
MiniMax-M2.7-highspeedMiniMax204,800131,072
MiniMax M2.1AWS Bedrock204,800131,072
GLM-4.7AWS Bedrock204,800131,072
MiniMax M2AWS Bedrock204,608128,000
Google Gemma 3 27B InstructAWS Bedrock202,7528,192
GLM-5AWS Bedrock202,752101,376
o3OpenAI200,000100,000
o4-miniOpenAI200,000100,000
Claude Haiku 4.5Anthropic200,0008,192
Sonar ProPerplexity200,000
Claude Opus 4.1 (latest)Anthropic200,00032,000
Claude Opus 4.5 (latest)Anthropic200,00064,000
GLM-4.7-FlashZ.ai (Zhipu AI)200,000131,072
GLM-4.7-FlashXZ.ai (Zhipu AI)200,000131,072
GLM-5.1Z.ai (Zhipu AI)200,000131,072
GLM-5-TurboZ.ai (Zhipu AI)200,000131,072
GLM-5V-TurboZ.ai (Zhipu AI)200,000131,072
o1OpenAI200,000100,000
o1-proOpenAI200,000100,000
o3-deep-researchOpenAI200,000100,000
o3-miniOpenAI200,000100,000
o3-proOpenAI200,000100,000
o4-mini-deep-researchOpenAI200,000100,000
GLM-4.7-FlashAWS Bedrock200,000131,072
MiniMax-M2MiniMax196,608128,000
MiniMax M2.5AWS Bedrock196,60898,304
DeepSeek V3.2DeepSeek164,0008,192
Hunyuan A13B InstructTencent131,0728,192
GLM-4.5Z.ai (Zhipu AI)131,07298,304
GLM-4.5-AirZ.ai (Zhipu AI)131,07298,304
Google Gemma 3 12BAWS Bedrock131,0728,192
QVQ MaxAlibaba131,0728,192
Qwen3 Coder NextAWS Bedrock131,07265,536
Qwen-VL MaxAlibaba131,0728,192
Qwen-VL PlusAlibaba131,0728,192
Qwen2.5 14B InstructAlibaba131,0728,192
Qwen2.5 32B InstructAlibaba131,0728,192
Qwen2.5 72B InstructAlibaba131,0728,192
Qwen2.5 7B InstructAlibaba131,0728,192
Qwen2.5-VL 72B InstructAlibaba131,0728,192
Qwen2.5-VL 7B InstructAlibaba131,0728,192
Qwen3 14BAlibaba131,0728,192
Qwen3 235B-A22BAlibaba131,07216,384
Qwen3 32BAlibaba131,07216,384
Qwen3 8BAlibaba131,0728,192
Qwen3-Next 80B-A3B InstructAlibaba131,07232,768
Qwen3-Next 80B-A3B (Thinking)Alibaba131,07232,768
Qwen3-VL 235B-A22BAlibaba131,07232,768
Qwen3-VL 30B-A3BAlibaba131,07232,768
QwQ PlusAlibaba131,0728,192
Mistral Medium 3.1Mistral AI131,0008,192
Mistral Small 4Mistral AI131,0008,192
GPT-5OpenAI128,000128,000
GPT-5 MiniOpenAI128,00064,000
GPT-4oOpenAI128,00016,384
GPT-4o miniOpenAI128,00016,384
Amazon Nova MicroAWS Bedrock128,0005,120
Llama 3.3 70BMeta128,0008,192
Command R+Cohere128,0004,096
Command RCohere128,0004,096
Command R7BCohere128,0004,096
SonarPerplexity128,000
Sonar Reasoning ProPerplexity128,000
Sonar Deep ResearchPerplexity128,000
GLM-4.6VZ.ai (Zhipu AI)128,00032,768
Gemma 3 4B ITAWS Bedrock128,0004,096
GPT-4 TurboOpenAI128,0004,096
GPT-5.3 Codex SparkOpenAI128,00032,000
Llama-3.3-70B-InstructMeta128,0004,096
Llama-4-Maverick-17B-128E-Instruct-FP8Meta128,0004,096
Magistral SmallMistral AI128,000128,000
Ministral 14B 3.0AWS Bedrock128,0004,096
Ministral 3 8BAWS Bedrock128,0004,096
Mistral NemoMistral AI128,000128,000
NVIDIA Nemotron Nano 12B v2 VL BF16AWS Bedrock128,0004,096
NVIDIA Nemotron Nano 3 30BAWS Bedrock128,0004,096
NVIDIA Nemotron Nano 9B v2AWS Bedrock128,0004,096
Open Mistral NemoMistral AI128,000128,000
gpt-oss-120bAWS Bedrock128,00016,384
gpt-oss-20bAWS Bedrock128,00016,384
Pixtral 12BMistral AI128,000128,000
Qwen3-Omni FlashAlibaba65,53616,384
GLM-4.5VZ.ai (Zhipu AI)64,00016,384
Mixtral 8x22BMistral AI64,00064,000
Qwen MaxAlibaba32,7688,192
Qwen-Omni TurboAlibaba32,7682,048
Qwen2.5-Omni 7BAlibaba32,7682,048
Step 1 (32K)StepFun32,76832,768
Mixtral 8x7BMistral AI32,00032,000
GPT-3.5-turboOpenAI16,3854,096
Phi 4Microsoft16,38416,384
Qwen-MT PlusAlibaba16,3848,192
Qwen-MT TurboAlibaba16,3848,192
Step 2 (16K)StepFun16,3848,192
GPT-4OpenAI8,1928,192
Qwen Plus Character (Japanese)Alibaba8,192512
Mistral 7BMistral AI8,0008,000

Then: Will It Fit? Context Checker

Results update automatically as you type.

Result
Yes — it fits in one call
53,000 tokens needed vs 128,000 tokens window
Document
50,000 tokens
Prompt / instructions
1,000 tokens
Reserved for output
2,000 tokens
Total needed
53,000 tokens
Context window
128,000 tokens
Docs of this size per call
2

Why this isn't trivial

The part people underestimate: few models support 200k tokens, and among those the input rate dominates because you are filling a huge window. In practice the biggest savings come from compressing context or using RAG so you need a smaller, cheaper window, so it is worth modelling before you commit.

How it's calculated

We estimate this by filtering models by context window, then ranking the qualifying ones by input cost at 200k tokens. Every figure uses the current provider prices baked into the site (reviewed daily), and you can override any input to match your own assumptions.

Frequently asked questions

Do I really need 200k tokens?+

Often not — RAG can retrieve just the relevant parts far more cheaply.

Why is filling a big window costly?+

You pay per input token, so 200k tokens is a large input bill every call.

Are these prices up to date?+

Yes. The model prices behind this calculator are refreshed and reviewed daily, so your estimate reflects current provider rates rather than a stale snapshot.

Related

Estimates for planning. Pricing data last reviewed 28 July 2026.