Claude API cost calculator

Anthropic bills Claude per token, with separate rates for input, output and the prompt cache. Caching is opt-in: you mark the shared prefix with cache_control, and writing it costs more than normal input, 1.25× the input price for the 5-minute cache and 2× for the 1-hour cache. Each later request that reuses the prefix reads it at 2.5%–10% of the input price, depending on the model, and every read restarts the cache lifetime.

The discounts stack. The Batch API takes 50% off input and 50% off output, and a cache read inside a batch costs 50% of the normal cache-read price, so batched jobs that share a prefix are the cheapest way to run Claude. Long prompts work the other way: long-context rates apply above 100K for Claude Haiku 5.5 (input 5×, output 5×); above 200K for Claude Sonnet 4.5 (input 2×, output 1.5×), and then to the whole request. The threshold counts every input token, cached reads and writes included, so caching a large document does not keep a request under it.

The table prices one workload on each model: 10,000 requests a day with 6,000 input and 500 output tokens per request. Columns: no caching; 70% of input read from the cache and 5% written to it; every request sent through the Batch API. For the same traffic Claude Fable 5.1 costs 100× as much as Claude Haiku 5.5.

ModelStandard
10,000 req/day · 6,000 in · 500 out
Calculate with your numbers →
With prompt caching
10,000 req/day · 6,000 in · 500 out · cache 70%
Calculate with your numbers →
Batch API
10,000 req/day · 6,000 in · 500 out · batch
Calculate with your numbers →
Claude Sonnet 5.5$5,100/mo$2,751/mo$2,550/mo
Claude Opus 5.5$10,200/mo$5,502/mo$5,100/mo
Claude Fable 5.1$25,500/mo$13,440/mo$12,750/mo
Claude Haiku 5.5$255/mo$144/mo$128/mo
Claude Haiku 4.5$2,550/mo$1,438/mo$1,275/mo

Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.

When does caching pay off for your traffic? Use the prompt caching calculator →

FAQ

Is the 1-hour cache worth it?

Only when the same prefix comes back at gaps longer than five minutes but within the hour, for example a support bot with sparse traffic. A 1-hour write costs more than a 5-minute write, so with steady traffic the 5-minute cache, refreshed by every hit, is cheaper. The prompt caching calculator compares both for your traffic.

Can I combine the Batch API with prompt caching on Claude?

Yes. Batch requests can read and write the cache, and cache prices are discounted inside a batch as well. Hits are less predictable because batch requests run in no fixed order, so send the requests that share a prefix in the same batch.

Why does my Claude bill jump on long prompts?

Some Claude models switch to long-context rates once one request's input passes a threshold. The whole request, cached tokens included, is then billed at the higher input and output rates. Keep prompts under the threshold with retrieval, or pick a model without a long-context tier.