LLM API cost calculator
Pick a preset, adjust your traffic, and every model's monthly cost updates instantly using its real cache, batch and long-context rates. Tick models to compare them, or share the link to reproduce the result.
| Compare | Model | Monthly cost | Savings | Per 1K requests | Input /1M | Output /1M | Cached input /1M | Provider | Notes |
|---|
Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.
Selling a subscription? Check your margin after AI API cost. Running a support bot? See the chatbot API cost at three traffic levels.
How the estimate works
Per request: uncached input × input price + cache reads × cache-read price + cache writes × cache-write price (normal input price for providers that cache automatically) + output × output price. Prompts above a model's long-context threshold switch the whole request to its long-context rates. Batch requests use published batch rates. Results are estimates: cache lifetime (TTL), reasoning tokens and tool calls change real bills.
FAQ
Does cached input cost less?
Yes. Providers bill input tokens read from the prompt cache at a fraction of the normal input price, often 10%. Some providers, such as Anthropic, also charge separately to write the cache. When a model has a published cache-write rate, this calculator includes it in the estimate.
What does the batch toggle do?
Batch share is the percentage of requests you send through a Batch API (asynchronous, usually within 24 hours). Those requests use each model's published batch input, output and cache rates. Models without a batch price keep the normal price and are marked "no batch".
Why is my real bill different?
Reasoning models bill hidden thinking tokens as output, long prompts can trigger higher long-context rates, and tool calls add tokens. Treat results as an estimate.