OpenAI API cost calculator
OpenAI caches repeated prompt prefixes automatically: there is nothing to mark in the request, and a cache hit is billed at 5%–10% of the input price on the models below. In our data GPT-5.4 mini, GPT-5.4 nano have no cache-write fee, so a miss costs the same as normal input, while GPT-6.1 Sol, GPT-6 Astra, GPT-6 Luna list a cache-write price of 1.25× the input rate. GPT-5.5 Pro has no cached-input price at all, so caching saves nothing on it.
The Batch API takes 50% off input and 50% off output in exchange for asynchronous results, and cached tokens inside a batch cost 50% of the normal cached rate. Long-context rates apply above 272K for GPT-6.1 Sol, GPT-6 Astra, GPT-5.5 Pro, GPT-6 Luna and 7 more (input 2×, output 1.5×), to the whole request. The biggest lever is the model tier: pro models spend more compute per answer and charge far more per token, mini and nano trade quality for price, and the flagship sits in between.
The table prices one workload on each model: 10,000 requests a day with 6,000 input and 500 output tokens per request. Columns: no caching; 70% of input read from the cache and 5% written to it; every request sent through the Batch API. For the same traffic GPT-5.5 Pro costs 318× as much as GPT-6 Luna.
| Model | Standard 10,000 req/day · 6,000 in · 500 out Calculate with your numbers → | With prompt caching 10,000 req/day · 6,000 in · 500 out · cache 70% Calculate with your numbers → | Batch API 10,000 req/day · 6,000 in · 500 out · batch Calculate with your numbers → |
|---|---|---|---|
| GPT-6.1 Sol | $5,100/mo | $2,751/mo | $2,550/mo |
| GPT-6 Astra | $25,500/mo | $14,385/mo | $12,750/mo |
| GPT-5.5 Pro | $81,000/mo | $81,000/mo no caching | $40,500/mo |
| GPT-6 Luna | $255/mo | $144/mo | $128/mo |
| GPT-5.4 mini | $2,025/mo | $1,174/mo | $1,012/mo |
| GPT-5.4 nano | $548/mo | $321/mo | $274/mo |
Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.
When does caching pay off for your traffic? Use the prompt caching calculator →
FAQ
Do I need to change my code to use OpenAI prompt caching?
No. Caching is automatic once a prompt is long enough. Put fixed content (instructions, tool definitions, examples) at the start and the variable part at the end so the prefix matches across requests.
When is a pro model worth it?
When a wrong answer costs more than the extra tokens: hard reasoning, high-stakes reviews, low volume. Pro models have no cached-input price in our data, so a high-volume app pays the full input price on every request.
Should I use mini or nano?
Nano suits classification, extraction and routing at very high volume; mini handles chat and simple reasoning. Test both on your own prompts: a cheaper model that needs retries or longer prompts can end up costing more.