Gemini API cost calculator

Gemini caches repeated prompt prefixes implicitly. There is no cache-write fee in our data: the first request pays normal input and later hits are billed at 10% of the input price. Caching only starts above a minimum prompt length (4,096 tokens where our data lists one), so short prompts never get the discount, and you cannot force a hit. Google also sells explicit caches with an hourly storage fee; those are not modelled here.

The Batch API takes 50% off input and 50% off output, but cached tokens inside a batch cost 50%–100% of the normal cached rate depending on the model, so Batch and caching do not always stack. Long-context rates apply above 200K for Gemini 3.1 Pro Preview and 2 more (input 2×, output 1.5×), to the whole request; models not named there keep one price at every prompt length.

The table prices one workload on each model: 10,000 requests a day with 6,000 input and 500 output tokens per request. Columns: no caching; 70% of input read from the cache and 5% written to it; every request sent through the Batch API. For the same traffic Gemini 3.1 Pro Preview costs 8× as much as Gemini 3.1 Flash Lite.

ModelStandard
10,000 req/day · 6,000 in · 500 out
Calculate with your numbers →
With prompt caching
10,000 req/day · 6,000 in · 500 out · cache 70%
Calculate with your numbers →
Batch API
10,000 req/day · 6,000 in · 500 out · batch
Calculate with your numbers →
Gemini 3.5 Flash$4,050/mo$2,349/mo$2,025/mo
Gemini 3.1 Pro Preview$5,400/mo$3,132/mo$2,700/mo
Gemini 3.8 Flash$1,912/mo$1,062/mo$956/mo
Gemini 3.5 Flash Lite$915/mo$575/mo$458/mo
Gemini 3.1 Flash Lite$675/mo$392/mo$338/mo

Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.

When does caching pay off for your traffic? Use the prompt caching calculator →

FAQ

Do I pay to write Gemini's implicit cache?

Not in our data: a cache miss is billed as normal input and hits get the cached rate. Keep the shared part at the start of the prompt and send similar requests close together to raise the hit rate.

Which Gemini models charge more for long prompts?

Only the models named in the long-context sentence above; all others keep one price at every prompt length in our data. Above the threshold the whole request uses the higher input and output rates.

Is Gemini Batch cheaper than caching?

For unique prompts, Batch saves more. For prompts that share a long prefix, caching saves more on input. Use both where you can, but check the cached rate inside batches, which is not discounted further on every model.