Best LLM for every task
The right model depends on the job. Each guide below picks the cheapest model that’s still smart enough for that task — ranked by community Arena score and price — so you don’t overpay for capability you can’t use.
The best value LLMs for writing and debugging code, ranked by coding Arena score and price. See the cheapest model that's still strong at code.
The best value LLMs for a support chatbot, ranked by Arena score and cost per month. High volume makes price the deciding factor.
The best value LLMs for summarizing text, ranked by Arena score and price. Summarization is input-heavy, so input pricing drives cost.
The best value LLMs for multi-step agents and tool use, ranked by agentic Arena score and cost. Reliability across many calls is what counts.
The best value LLMs for math and multi-step reasoning, ranked by reasoning Arena score and price.
The best value LLMs for long documents and large context windows (200K–1M tokens), ranked by Arena score and price.
The best value multimodal LLMs for image understanding, ranked by Arena score and price. Only vision-capable models are shown.
The cheapest LLMs that are still genuinely usable, ranked by price with a minimum quality bar so you don't ship something too weak.
The best value LLMs for high-volume classification and tagging, ranked by cost with a sane quality bar. Tiny outputs make price per call tiny — volume makes it add up.
The best value LLMs for RAG pipelines, ranked by Arena score and price. Retrieval does the heavy lifting, so a mid-tier model often suffices.
The best value LLMs for extracting structured data (JSON, fields, tables) from text, ranked by Arena score and price.
The best value LLMs for content and copywriting, ranked by Arena score and price. Fluency is table stakes, so value comes from cost.
Best model for your coding tool
Editors and agents like Cursor, Copilot and Aider let you choose the underlying model. These guides rank the models each tool supports by coding value, so you spend your request pool (or your API bill) where it counts.
Which model to pick inside Cursor: the coding models Cursor supports, ranked by coding Arena score and cost so you get the best value from your agent-request pool.
Which model to select in GitHub Copilot: the frontier models Copilot exposes for chat and agent mode, ranked by coding Arena score and price to make premium requests count.
Which model to use in Windsurf's Cascade agent: the frontier models it supports, ranked by coding Arena score and cost so your monthly credits go further.
Aider is bring-your-own-key, so the model IS the cost. The best value models for Aider's terminal pair-programming, ranked by coding Arena score and price per 1M tokens.
Cline is an open-source VS Code agent that uses your own API key. The best value models for Cline's autonomous edits, ranked by coding Arena score and token price.
Continue is an open-source IDE assistant you wire to your own models. The best value models for Continue's chat, edit and autocomplete, ranked by coding Arena and price.
Bolt.new generates and runs full-stack apps in the browser. The best value models for Bolt-style app generation, ranked by coding Arena score and token cost.
Intelligence scores are community Arena ratings from LMArena / arena.ai, used under CC BY 4.0. Snapshot last refreshed 28 July 2026. Scores are a relative signal, not an absolute measure of capability. A Measured score was voted on directly; an Estimated score is inherited from an identical base model (a regional/creator-prefixed hosting duplicate). See our methodology.