Estimate Your Support Chatbot’s Monthly API Cost

A support chatbot’s API budget starts with how often people use it and how much context each reply needs. Use this worked example to turn those assumptions into a monthly estimate.

A worked monthly estimate

A planning example for 1,000 monthly users. Start with the workload, then inspect the API bill it produces.

What we’re pricing

1,000 people using support chat each month

Messages per person / month
20
Model calls per message
1
Input tokens per call
2,000
Output tokens per call
500

These are illustrative workloads, not measured customer bills. Activity counts are monthly; token counts are per model call.

Adapt this estimate to my app

No account required. The link preserves these inputs and data versions.

ByeTokens

Example estimate
Support chatbot
1,000 users / month

Monthly API cost$9.00

  • AI responses
    GPT-6 Luna$9.00
    50,000,000 tokens20,000 requestsGood fit: support chat

Total$9.00/mo

Catalog 2026-09-26-7e5c5fac
Thank you for not counting tokens.

Tariffs checked 2026-09-26T12:00:00.000Z · Snapshot generated 2026-09-26T12:00:00.000Z

Lowest-cost complete compatible stack with fresh catalog prices. No provider locks, extra capability filters or recurring free allowances. This is not a quality recommendation.

How we calculated this

Official pricing sources, billable quantities and formulas for the 1,000-user example.

LLM
gpt 6 luna, plan openai-gpt-6-luna-standard
$9.00
Billable quantity
20,000 requests
Fixed fee
$0.00
Usage
$9.00
Formula
input tokens + output tokens
Fit for your tasksGood fit
Support chat80 of 100, likely 75-86
  • LMArena text, instruction following35% weight, measured as gpt-6-luna-max, Sep 25, 202682
  • LMArena text, multi-turn35% weight, measured as gpt-6-luna-max, Sep 25, 202678
  • LMArena text, overall30% weight, measured as gpt-6-luna-max, Sep 25, 202679

Fit scores are leaderboard percentiles: 90 means better than 90% of the models listed. Quality data: LMArena and Epoch AI, CC BY 4.0. Exact engine totals are summed before display rounding. Taxes, infrastructure and retry rates are excluded. Enabled affiliate links are marked individually and do not affect rankings. Disclosure

How the budget changes with usage

The same per-person workload at three audience sizes. Each row is recalculated, so providers and plans can change with volume.

Monthly users Model calls Monthly API estimate Try these inputs
1002,000$0.90Calculate for 100 users →
1,00020,000$9.00Calculate for 1,000 users →
10,000200,000$90.00Calculate for 10,000 users →

Catalog 2026-09-26-7e5c5fac. Calculated when this page loads; an archived scenario may differ from current prices.

Start with messages, not registered accounts

A thousand registered accounts do not tell you how many model calls you will pay for. Some people may never open support; others may send several messages in one conversation. Estimate the number of people using the feature each month, then the messages each person sends.

In the worked example above, every listed monthly user sends the same average number of messages. If your app has a larger audience but only a fraction uses the assistant, enter the people who use it, or adjust the average across your whole audience. Do not apply that fraction twice.

A message means one user action that triggers the configured number of model calls. It is not a resolved ticket or an entire conversation. The example makes one call per message. A workflow that classifies a request, drafts a response and checks it may need several calls; account for those before relying on the estimate.

What is inside each model call?

Input includes the instructions, the user’s message and any context you send with it. That context might contain previous turns or passages from a knowledge base. Output is the generated response. The table above makes both quantities explicit so you can replace them with a representative sample from your app.

The example uses a fixed average per call. It does not simulate a conversation whose history grows at every turn. If long sessions are common, measure or estimate that larger average. Output that looks short to the user may also require billable reasoning; check your chosen provider’s accounting and include it in the workload where applicable.

For each workload, the engine prices input and output separately using the selected model’s applicable rate band. The calculation trace shows the actual formula. It does not average a small prompt with a large one before choosing the context band.

Read the three usage levels

The usage table keeps the per-person assumptions fixed and changes the monthly audience. Each row is a new calculation against the same catalog version. This shows the effect of volume without quietly changing the task.

The selected option is the lowest-cost complete compatible stack with fresh prices in the supported catalog. Low cost does not prove that a model will answer support questions correctly. Test candidate models on representative tickets and your escalation rules. In the calculator, Quality fit is a separate evidence-based view; it is not a guarantee for your own support data.

To inspect provider coverage, start with OpenAI, Anthropic or Google. Their pages show the specific offers we can price, rather than every product those companies sell.

What this estimate leaves out

This support example prices the model calls. It does not include your application hosting, storage, vector database, embeddings, human support staff or engineering time. Supplying retrieved text as model input does not account for the cost of retrieving or indexing that text.

Retries and extra calls are not added automatically. Caching, batch discounts, negotiated rates and taxes are outside this example. Recurring free allowances are switched off. Check the methodology before treating the estimate as a project budget.

If your assistant searches the live web or reads external pages, use the research-agent guide to understand those separate API charges. You can combine Chat and Research in the calculator, but make sure you do not count the same model call twice.

Turn the example into your budget

Open a row from the table and change the audience and activity first. Then inspect the token assumptions and provider constraints. The link preserves the example’s inputs and data versions so you can see what changed. A saved price may age; review its date before choosing a service.

Use the estimate to narrow the options and plan a small real workload test. Compare the resulting usage with your assumptions, then update the budget. That is a stronger basis for a launch decision than assuming every conversation will behave like the demo.