Prompt caching pricing & savings

When many requests start with the same long prefix (system prompt, tool definitions, a shared document), providers can cache it and bill those tokens at a much lower cached-input rate. Chatbots and agent loops benefit most because they resend the same context every turn.

Estimate your monthly bill with this discount →

ModelNormal input /1MCached input /1MSaving
DeepSeek V4 Flash$0.30$0.00698%
Claude Fable 5.1$10.00$0.2598%
Claude Mythos 5.1$10.00$0.2598%
DeepSeek V4 Pro$1.32$0.04497%
Claude Opus 5.5$4.00$0.2095%
GPT-6.1 Sol$2.00$0.1095%
Claude Sonnet 4.5$3.00$0.3090%
Claude Sonnet 4.6$3.00$0.3090%
Gemini 2.5 Flash Lite$0.10$0.0190%
Gemini 3.5 Flash$1.50$0.1590%
Gemini 3.6 Flash$0.75$0.07590%
Gemini 3.7 Flash$0.75$0.07590%
Gemini 3.8 Flash$0.75$0.07590%
Gemini Flash (latest)$0.75$0.07590%
Devstral Medium (latest)$0.40$0.0490%
Devstral Small (latest)$0.10$0.0190%
Magistral Medium (latest)$1.50$0.1590%
Ministral 14b (latest)$0.20$0.0290%
Ministral 3b (latest)$0.10$0.0190%
Mistral Medium 3$1.50$0.1590%
Mistral Medium 3.5$1.50$0.1590%
GPT-5 nano$0.05$0.00590%
GPT-5.2$1.75$0.17590%
GPT-5.4 mini$0.75$0.07590%
GPT-5.4 nano$0.20$0.0290%
GPT-5.6 Luna$0.20$0.0290%
GPT-6 Luna$0.10$0.0190%
Claude Fable 5$10.00$1.0090%
Claude Haiku 4.5$1.00$0.1090%
Claude Mythos 5$10.00$1.0090%
Claude Opus 4.5$5.00$0.5090%
Claude Opus 4.6$5.00$0.5090%
Claude Opus 4.7$5.00$0.5090%
Claude Opus 4.8$5.00$0.5090%
Claude Opus 5$5.00$0.5090%
Claude Sonnet 5$2.00$0.2090%
Claude Sonnet 5.5$2.00$0.2090%
Gemini 2.5 Flash$0.30$0.0390%
Gemini 2.5 Pro$1.25$0.12590%
Gemini 3 Flash Preview$0.50$0.0590%
Gemini 3.1 Flash Lite$0.25$0.02590%
Gemini 3.1 Pro Preview$2.00$0.2090%
Gemini 3.5 Flash Lite$0.30$0.0390%
Gemini Flash Lite (latest)$0.30$0.0390%
Gemini Pro (latest)$2.00$0.2090%
Codestral (latest)$0.30$0.0390%
Magistral Small (latest)$0.15$0.01590%
Ministral 8b (latest)$0.15$0.01590%
Mistral Large (latest)$0.50$0.0590%
Mistral Large 4$0.68$0.06890%
Mistral Small (latest)$0.15$0.01590%
GPT-5$1.25$0.12590%
GPT-5 mini$0.25$0.02590%
GPT-5.1$1.25$0.12590%
GPT-5.4$2.50$0.2590%
GPT-5.5$5.00$0.5090%
GPT-5.6 Sol$4.00$0.4090%
GPT-5.6 Terra$2.00$0.2090%
GPT-6 Astra$10.00$1.0090%
GPT-6 Sol$2.00$0.2090%
Grok 4.5$2.00$0.3085%
Grok 4.20 Non Reasoning$1.25$0.2084%
Grok 4.20 Reasoning$1.25$0.2084%
Grok 4.3$1.25$0.2084%
Grok Build 0.1$1.00$0.2080%
Grok Code Fast 1$1.00$0.2080%
GPT-4.1$2.00$0.5075%
GPT-4.1 mini$0.40$0.1075%
o3$2.00$0.5075%
Grok 4.6$2.00$0.5075%
Grok 4.7$2.00$0.5075%
GPT-4o$2.50$1.2550%
GPT-4o mini$0.15$0.07550%

FAQ

Does writing to the cache cost extra?

Anthropic charges a premium (shown as cache write on model pages) the first time a prefix is cached; OpenAI and Google cache automatically without a write fee in most cases.

How long does a cached prompt last?

Typically minutes (often around 5) unless refreshed by new requests; some providers sell longer cache lifetimes.

What cache hit rate is realistic?

Agent loops and multi-turn chats often reach 50–80% of input tokens; one-off requests with unique documents get close to 0%.