1 call per message, 2,000 input and 500 output tokens
Every total should be traceable.
A plain-language guide and technical reference for workload normalization, pricing formulas, recommendations, evidence, and known exclusions.
Start with a worked example: support chatbot or research agent. These are API estimates under explicit assumptions, not complete infrastructure budgets.
- Catalog
- 2026-09-26-7e5c5fac
- Providers
- 11
- Models and endpoints
- 20
- Reviewed
- Sep 24, 2026
On this page
Every number gets a stamp.
- Verified
Price read from the official page within the last 14 days.
- Expired
Price older than 14 days. Still shown, never passed off as fresh.
- Archived
A saved receipt shown with the catalog it was made with.
- No data
Not enough independent evidence to score it, so no rating.
- Strong fit
Top 10% of independent leaderboards for your task.
- Good fit
Top 30% for your task. Balanced picks the cheapest stack at this level or better.
- Limited fit
Below the top 30%. Fine for some jobs, but not what Balanced recommends.
- Catalog only
Provider listed, but its billing does not fit a fair calculation yet.
How fit is scored
Rendered from the same task profiles the calculator uses.
Each signal is the model’s percentile on an independent public leaderboard. Strong is the top 10% of the weighted blend, good the top 30%. A task needs signals carrying 60% of its weight and nothing older than 90 days.
| Task | Signals and weights |
|---|---|
| General chat | LMArena text, overall 60%, LMArena text, instruction following 20%, LMArena text, multi-turn 20%
|
| Support chat | LMArena text, instruction following 35%, LMArena text, multi-turn 35%, LMArena text, overall 30%
|
| Coding help | LMArena text, coding 50%, LMArena WebDev 50%
|
| Document work | LMArena text, long queries 50%, LMArena documents 30%, LMArena text, instruction following 20%
|
| Research agent | LMArena text, hard prompts 35%, LMArena agents 25%, SimpleQA Verified (Epoch AI) 20%, GPQA Diamond (Epoch AI) 20%
|
| Image generation | LMArena text-to-image, overall 70%, LMArena text-to-image, photorealistic 15%, LMArena text-to-image, commercial design 15%
|
| Video generation | LMArena text-to-video 60%, LMArena image-to-video 40%
|
Default scenario assumptions
Rendered from the same typed defaults used by the calculator.
3 LLM calls, 2 searches and 3 pages per task
1024 × 1024 px
5 seconds, 720p, no audio
ByeTokens answers a narrow question: what would the catalogued API services cost each month for a defined application workload? It translates product activity into billable units, calculates every compatible stack with decimal arithmetic, and explains which plan and formula produced each total.
The result is an estimate, not a provider quote and not a complete infrastructure budget. It is designed to be reproducible: the same scenario, catalog version, evidence version, and calculation rules should produce the same answer.
The short version
You describe the features in an application and their expected monthly usage. The calculator converts those inputs into language-model calls, searches, scraped pages, images, and video clips. Each category is priced with the exact billing rule stored for a provider offer. Category totals are added to form a stack total.
Cheapest selects the lowest exact cost among compatible stacks. Fastest and Balanced require recent evidence that can be compared fairly. If that evidence does not exist, those recommendations remain unavailable while the cost table continues to work.
Prices come from public provider sources recorded in a versioned catalog. Every plan includes source links and a verification timestamp. Historical catalog versions remain available so a saved or shared scenario can be reproduced after current prices change.
What the estimate includes
The estimate includes charges represented by the selected offers and plans: token usage, metered operations, package fees and overage, hard-cap subscriptions, image outputs, and video outputs. Recurring free allowances are ignored by default and can only be included when the plan explicitly marks them as recurring.
The estimate excludes taxes, application hosting, databases, storage, observability, network delivery, engineering work, and unmodelled retry behavior. It also excludes enterprise agreements, annual commitments, batch pricing, caching discounts, and optional endpoint features unless a dedicated offer states otherwise.
Unknown price never becomes zero. If an active category has no compatible offer with a complete billing rule, that stack is unavailable.
Current scenario assumptions
The table above this article is generated from the same shared defaults used by the calculator. It is not copied into Markdown. When a product assumption changes in code, the methodology page changes with it.
Scenario mode is intended for early product planning. Exact mode is intended for measured or forecast monthly totals. Switching to exact mode copies the currently derived workload, after which the exact fields become the only calculation source.
From product activity to workload
Chat messages are calculated as the ceiling of monthly users multiplied by messages per user. Language-model calls then multiply that message count by calls per message.
Research tasks follow the same pattern. Each task can create several language-model calls, searches, and page retrievals. Chat and research language work remain separate rows because their token profiles may differ. One selected language-model offer serves both rows, but context-band selection and pricing happen before the rows are summed.
Image demand is the ceiling of users multiplied by images per user. Video demand uses users multiplied by clips per user. Width, height, duration, resolution, and audio remain properties of the output profile; they are not multiplied by growth controls.
Derived scenario counts are rounded upward once at the scenario boundary. Usage per user can retain up to three decimal places. Exact monthly counts must be integers.
Supported billing formulas
Language-model tokens
For each language workload, input cost equals total input tokens divided by one million and multiplied by the applicable input rate. Output cost uses the same calculation with output tokens and the output rate. The workload total is input plus output.
When rates change above a context threshold, the engine selects the band independently for each workload. It does not average chat and research tokens before choosing a band.
Metered operations
A metered plan begins with a fixed fee, subtracts included units from demand, and prices the remaining units in the plan’s rate quantity. The payable quantity cannot fall below zero.
Packages with overage
A package plan includes a fixed allowance. Demand above that allowance is divided by the package size and rounded upward, then multiplied by the package price. This creates step changes at quota boundaries.
Hard caps
A hard-cap plan charges its fixed fee while demand remains inside the included quantity. If demand exceeds the verified cap and no overage rule exists, the plan is unavailable. The calculator does not invent an overage price.
Image outputs
An image offer can charge per image or per megapixel. Megapixels equal width multiplied by height divided by one million. When a megapixel billing step exists, the billable area rounds upward to that step before the rate is applied.
Video outputs
A video offer can charge per fixed-profile clip or per output second. Per-second rules can define minimum and step durations. The requested duration is first raised to the minimum, then rounded upward to the billing step. Resolution and audio must match the verified offer profile.
A synthetic example
Suppose a fictional plan charges a $10 monthly fee, includes 1,000 operations, and sells additional blocks of 500 operations for $4. A workload of 1,001 operations needs one overage block, so the total is $14. A workload of 1,500 operations is also $14. At 1,501 operations the total becomes $18.
This example exists only to explain package rounding. It is synthetic test data, not a real provider price.
Compatibility before price
An offer is considered only when it supports the requested category and profile. Language offers can be filtered by minimum context, tool calling, structured output, and token limits. Media offers must match dimensions or megapixel handling, duration, resolution, and audio requirements.
Unknown capability data fails closed when the user requires that capability. A locked component does not bypass compatibility. If a lock is incompatible, the result explains the exclusion instead of silently replacing the selection.
Zero demand disables a category and its subscription fee. This prevents an unused feature from adding a monthly plan charge.
Plans, stacks, and baseline
The engine evaluates eligible plans for every active offer and selects the cheapest valid plan for the workload. It then creates compatible combinations across active categories. Candidate stacks are sorted by exact monthly cost and stable ID.
A baseline is optional. When present, it uses the same workload, catalog version, and calculation rules as candidates. Difference, percentage, annualized difference, and cost per user are derived from exact totals. A zero baseline never produces an infinite or misleading percentage.
Displayed currency values are rounded for readability after calculation. CSV exports retain the exact decimal strings used by the engine.
Cheapest, Fastest, and Balanced
Cheapest selects the lowest-cost compatible candidate and does not depend on affiliate configuration.
Balanced selects the cheapest stack whose every scored component is at least a good fit for the tasks in your workload. Fit comes from independent public leaderboards, LMArena and Epoch AI, never from vendor announcements. For each leaderboard signal, a model gets its percentile: the share of listed models it beats. A task combines several signals with published weights. Strong means the top 10% on that blend, good the top 30%, and anything lower is limited. When the uncertainty range reaches well below a tier, the lower tier is shown.
One language model serves both chat and research, so it is judged by the weaker of the two tasks. A task needs signals carrying at least 60% of its weight, and evidence older than 90 days is ignored. Without that, the component shows no data rather than an estimate. Search and scraping services have no independent comparable benchmark and are not scored.
Fastest requires a recent speed-evidence cohort with compatible model, host, workload, and methodology. No speed source is published yet, so Fastest says that there is not enough comparable data.
Ranges and growth
The answer-length range reruns the engine with defined short and long output-token assumptions. It does not apply a percentage multiplier to the current bill.
The growth chart reruns the complete calculation at each user level. Counts, plan selection, quotas, hard caps, and compatibility are evaluated again. This is why a tenfold increase in users can produce more or less than a tenfold increase in cost.
Data freshness and reproducibility
The production catalog stores source URLs, access times, effective dates when known, licenses, and notes. Catalog and evidence versions are content-derived and immutable in history. Current data can advance without changing an older saved estimate.
Freshness affects status and release readiness. A stale price can still be displayed with a warning, but it cannot silently appear freshly verified. Missing or disputed data remains unavailable until a source-backed entry replaces it.
Use the pricing JSON or CSV links on this page to inspect the current public dataset. Use calculation details in the calculator to connect a monthly total to its plan, quantity, formula, verification date, and source.
Inspect the inputs
The same snapshot and source records the calculator reads.
Compare your own usage
CSV comparisons preserve your reported input and output token volumes and apply current standard uncached API rates to both the source model and alternatives. Cached input is included once at the ordinary input rate; reasoning tokens are already part of output. Positive rate differences mean lower estimated cost, not historical savings. Aggregate usage cannot establish request lengths, context compatibility or equivalent quality. Where tariffs have request-length bands, we show the minimum and maximum across those bands. Only active, effective standard plans verified within 14 days are compared. Tools, Batch, Scale Tier, fine-tuning, non-text usage, taxes and infrastructure are excluded.