Prompt caching break-even calculator
Caching a long shared prefix is not always cheaper. Providers like Anthropic charge extra to write the cache (1.25× input for 5 minutes, 2× for 1 hour), and the cache expires when requests are too far apart. Enter your prefix size and traffic to see which option is cheapest.
| Option | Cache hit chance | Per request | Monthly | Pays off when |
|---|
How the calculation works
Requests are assumed to arrive at random over your active hours. A request reads the cache when the previous one came within the cache lifetime (each read refreshes it), so the hit chance is 1 − e^(−requests per minute × TTL in minutes). Misses pay the write price. Prompts above a model's long-context threshold use its long-context rates. Automatic caching (OpenAI, Google and others without a write price) is modelled with a 5-minute lifetime.
FAQ
Why can prompt caching make my bill higher?
Writing to the cache costs more than normal input on Anthropic models. If requests arrive further apart than the cache lifetime, most requests write again and few read, so you pay the premium without the discount.
Should I use the 5-minute or the 1-hour cache?
The 1-hour cache costs more to write but stays warm through gaps of up to an hour. It usually wins when requests come every few minutes to tens of minutes; with steady traffic every minute or more often, the 5-minute cache is cheaper.