Prompt caching pricing & savings
When many requests start with the same long prefix (system prompt, tool definitions, a shared document), providers can cache it and bill those tokens at a much lower cached-input rate. Chatbots and agent loops benefit most because they resend the same context every turn.
Estimate your monthly bill with this discount →
FAQ
Does writing to the cache cost extra?
Anthropic charges a premium (shown as cache write on model pages) the first time a prefix is cached; OpenAI and Google cache automatically without a write fee in most cases.
How long does a cached prompt last?
Typically minutes (often around 5) unless refreshed by new requests; some providers sell longer cache lifetimes.
What cache hit rate is realistic?
Agent loops and multi-turn chats often reach 50–80% of input tokens; one-off requests with unique documents get close to 0%.