Should I self-host an LLM or keep using an API?
Is it worth it? To decide between self-hosting an LLM and using an API, you have to weigh what the AI costs against what it saves or earns. This page helps infra teams weighing self-hosting put both sides on the same page and see the payback, not just the invoice.
Results update automatically as you type.
- API cost / request
- $0.000645
- GPU host / month
- $1,500.00
- Break-even volume
- 2,325,581 req/mo
Why this isn't trivial
The part people underestimate: self-hosting adds fixed GPU and ops cost that only pays off above a utilization threshold you may not have hit yet. In practice the biggest savings come from self-hosting only once utilization clears the break-even, so it is worth modelling before you commit.
How it's calculated
We estimate this by comparing monthly API spend against amortized GPU, hosting and ops cost for an equivalent model. Every figure uses the current provider prices baked into the site (reviewed daily), and you can override any input to match your own assumptions.
Frequently asked questions
What's the break-even?+
Where amortized GPU cost per request drops below the API's per-request price.
What's easy to forget?+
Ops, scaling, reliability and idle-GPU cost — include them all.
Are these prices up to date?+
Yes. The model prices behind this calculator are refreshed and reviewed daily, so your estimate reflects current provider rates rather than a stale snapshot.
Related
Estimates for planning. Pricing data last reviewed 28 July 2026.