RAG app API cost
In a RAG app almost every token is input: each question carries the retrieved chunks, so in our example a request has 8,000 input tokens but only 600 output tokens, at 2,000 questions a day. Only the stable part (instructions and often-retrieved passages) can be cached; we assume 30% read from the cache. Turning your documents into vectors is a separate, mostly one-off cost: see embedding prices.
Sending whole documents instead of top-k chunks changes the bill completely. With 300,000-token prompts, Claude Haiku 5.5 goes from $66.00 to $9,090 a month, because above 100K tokens the whole request is billed at its long-context rate; caching 80% of the document, as in a chat-with-your-files app where users ask several questions about the same file, brings it to $2,722. With whole documents the cheapest model here becomes MiMo V2.6 Flash at $2,530.
The table shows the cheapest current models for each scenario; models whose context window is too small are left out of the whole-document columns. Thresholds per model: long-context pricing.
| Model | Top-k retrieval 2,000 req/day · 8,000 in · 600 out Calculate with your numbers → | Top-k + caching 2,000 req/day · 8,000 in · 600 out · cache 30% Calculate with your numbers → | Whole document 2,000 req/day · 300,000 in · 600 out Calculate with your numbers → | Whole document + caching 2,000 req/day · 300,000 in · 600 out · cache 80% Calculate with your numbers → |
|---|---|---|---|---|
| Claude Haiku 5.5 | $66.00/mo | $53.64/mo | $9,090/mo long-context rate | $2,722/mo long-context rate |
| GPT-6 Luna | $66.00/mo | $53.64/mo | $3,627/mo long-context rate | $1,080/mo long-context rate |
| MiMo V2.6 Flash | $77.28/mo | $57.52/mo | $2,530/mo | $554/mo |
| Qwen3.8 Flash | $88.92/mo | $70.82/mo | $2,717/mo | $832/mo |
| GLM-5.3 Flash | $90.00/mo | $72.72/mo | $2,718/mo | $990/mo |
Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.
FAQ
How many chunks should I retrieve?
As few as keep answers correct. Every extra chunk is billed on every question, so four times the chunks means roughly four times the retrieved part of the input. Measure answer quality at a few values of k before settling.
Is a long-context model cheaper than RAG?
Rarely at scale. Sending a whole document costs far more input per question than a few retrieved chunks, and some models charge a higher rate above a size threshold. It can make sense at low traffic, or when the same document is cached and questioned many times.