AI chatbot API cost per month

A chat assistant resends the system prompt and the conversation so far with every message, so input piles up while each answer stays short. We assume 4 messages per conversation, each a request with 1,200 input and 400 output tokens, where 60% of the input is a cached prefix (instructions, help-center snippets, earlier turns) and 5% is written to the cache. People are waiting for the answer, so the Batch API is not an option.

Cost grows in a straight line with traffic. At 100,000 conversations a day, MiMo V2.6 Flash comes to $2,175 a month, about $0.72 per 1,000 conversations: that per-conversation figure is the one to compare with your cost per support ticket or your subscription price. A longer system prompt or more retrieved articles adds input to every message, so it raises the bill on every turn, not once per conversation.

The table lists the cheapest current models for this workload. Change any number in the calculator, or check your margin per subscriber.

Model1,000 conversations/day
4,000 req/day · 1,200 in · 400 out · cache 60%
Calculate with your numbers →
10,000 conversations/day
40,000 req/day · 1,200 in · 400 out · cache 60%
Calculate with your numbers →
100,000 conversations/day
400,000 req/day · 1,200 in · 400 out · cache 60%
Calculate with your numbers →
MiMo V2.6 Flash$21.75/mo$217/mo$2,175/mo
Claude Haiku 5.5$30.80/mo$308/mo$3,080/mo
GPT-6 Luna$30.80/mo$308/mo$3,080/mo
Qwen3.8 Flash$32.94/mo$329/mo$3,294/mo
GLM-5.3 Flash$35.23/mo$352/mo$3,523/mo
Mistral Small (latest)$38.74/mo$387/mo$3,874/mo

Estimates from published per-token rates; check official pricing before budgeting. Data synced from the LiteLLM price list 2026-10-08.

FAQ

How many tokens is one chat message?

A short question and answer is a few hundred tokens, but the request also carries the system prompt, retrieved help articles and earlier turns. Log the usage your API returns for a day of real traffic and put the averages into the calculator.

How do I cut chatbot API cost?

Keep the system prompt and tool definitions fixed and at the start so they cache, trim or summarize old turns, and route easy questions to a small model while escalating only hard ones to a larger model.