AI Calculator Pro

glm-5-2-fp8

Z.ai (Zhipu AI) · model · context window 393,216 tokens.

Pricing (per 1M tokens)

Input$0.8500
Output$3.10

Specifications

Context window393,216 tokens
Max output131,072 tokens
Modalitiestext
VisionNo
Tool useYes
ReasoningNo
Prompt cachingNo

Where else you can run glm-5-2-fp8

The same model is resold by 3 providers we track. The typical one charges 1.1× what the cheapest does — $1.50 against $1.41 per 1M tokens, blended 3:1. Routing the same workload to a different host is often a bigger saving than switching model.

ProviderInput / 1MOutput / 1MBlended
Vultr$0.8500$3.10$1.41
Gmicloud$0.9790$3.08$1.50
Ambient$1.20$4.20$1.95

Third-party listings via models.dev, not rates we verify with each host the way we do first-party pricing. Treat them as a shortlist to check, not a quote — and confirm quantisation, context limits and rate limits before switching, since a cheaper host is not always serving the same thing.

Estimate cost with glm-5-2-fp8

glm-5-2-fp8 is pre-selected below. Enter your token counts and volume to see the cost per request, per day and per month — then switch models to compare.

Results update automatically as you type.

Result
$0.00240 per request
~$72.00/month at 1,000 requests/day
Input cost
$0.000850
Output cost
$0.00155
Per request
$0.00240
Per day
$2.40
Per month
$72.00

Source: community-hosted (cheapest of 3 providers, vultr via models.dev). Prices are re-checked daily against the provider’s official pricing; this rate was last changed on 27 August 2026. A recent date here just means the price is still current, not that it moved.