Claude API cost calculator
Anthropic bills Claude per token, with separate rates for input, output and the prompt cache. Caching is opt-in: you mark the shared prefix with cache_control, and writing it costs more than normal input, 1.25× the input price for the 5-minute cache and 2× for the 1-hour cache. Each later request that reuses the prefix reads it at 2.5%–10% of the input price, depending on the model, and every read restarts the cache lifetime.
The discounts stack. The Batch API takes 50% off input and 50% off output, and a cache read inside a batch costs 50% of the normal cache-read price, so batched jobs that share a prefix are the cheapest way to run Claude. Long prompts work the other way: long-context rates apply above 100K for Claude Haiku 5.5 (input 5×, output 5×); above 200K for Claude Sonnet 4.5 (input 2×, output 1.5×), and then to the whole request. The threshold counts every input token, cached reads and writes included, so caching a large document does not keep a request under it.
The table prices one workload on each model: 10,000 requests a day with 6,000 input and 500 output tokens per request. Columns: no caching; 70% of input read from the cache and 5% written to it; every request sent through the Batch API. For the same traffic Claude Fable 5.1 costs 100× as much as Claude Haiku 5.5.
| Model | Standard 10,000 req/day · 6,000 in · 500 out Calculate with your numbers → | With prompt caching 10,000 req/day · 6,000 in · 500 out · cache 70% Calculate with your numbers → | Batch API 10,000 req/day · 6,000 in · 500 out · batch Calculate with your numbers → |
|---|---|---|---|
| Claude Sonnet 5.5 | $5,100/mo | $2,751/mo | $2,550/mo |
| Claude Opus 5.5 | $10,200/mo | $5,502/mo | $5,100/mo |
| Claude Fable 5.1 | $25,500/mo | $13,440/mo | $12,750/mo |
| Claude Haiku 5.5 | $255/mo | $144/mo | $128/mo |
| Claude Haiku 4.5 | $2,550/mo | $1,438/mo | $1,275/mo |
Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.
When does caching pay off for your traffic? Use the prompt caching calculator →
FAQ
Is the 1-hour cache worth it?
Only when the same prefix comes back at gaps longer than five minutes but within the hour, for example a support bot with sparse traffic. A 1-hour write costs more than a 5-minute write, so with steady traffic the 5-minute cache, refreshed by every hit, is cheaper. The prompt caching calculator compares both for your traffic.
Can I combine the Batch API with prompt caching on Claude?
Yes. Batch requests can read and write the cache, and cache prices are discounted inside a batch as well. Hits are less predictable because batch requests run in no fixed order, so send the requests that share a prefix in the same batch.
Why does my Claude bill jump on long prompts?
Some Claude models switch to long-context rates once one request's input passes a threshold. The whole request, cached tokens included, is then billed at the higher input and output rates. Keep prompts under the threshold with retrieval, or pick a model without a long-context tier.