Research Agent API Costs — LLM, Search and Scraping

A research feature can pay for model calls, web searches and page retrieval in the same task. This worked budget separates those charges and shows how the monthly estimate changes with usage.

A worked monthly estimate

A planning example for 1,000 monthly users. Start with the workload, then inspect the API bill it produces.

What we’re pricing

1,000 people using research each month

Tasks per person / month
5
Model calls per task
3
Input tokens per call
2,000
Output tokens per call
500
Searches per task
2
Pages read per task
3

These are illustrative workloads, not measured customer bills. Activity counts are monthly; token counts are per model call.

Adapt this estimate to my app

No account required. The link preserves these inputs and data versions.

ByeTokens

Example estimate
Research agent
1,000 users / month

Monthly API cost$125.75

  • AI responses
    GPT-6 Luna$6.75
    37,500,000 tokens15,000 requestsFit for research agent: not enough data yet
  • Web search
    Exa Search$70.00
    10,000 searches
  • Reading web pages
    ScraperAPI Basic Page$49.00
    15,000 pages

Total$125.75/mo

Catalog 2026-09-26-7e5c5fac
Thank you for not counting tokens.

Tariffs checked 2026-09-24T12:00:00.000Z · Snapshot generated 2026-09-26T12:00:00.000Z

Lowest-cost complete compatible stack with fresh catalog prices. No provider locks, extra capability filters or recurring free allowances. This is not a quality recommendation.

How we calculated this

Official pricing sources, billable quantities and formulas for the 1,000-user example.

LLM
gpt 6 luna, plan openai-gpt-6-luna-standard
$6.75
Billable quantity
15,000 requests
Fixed fee
$0.00
Usage
$6.75
Formula
input tokens + output tokens
Fit for your tasksNo data
Research agentNot enough data, 35% of signals
  • LMArena text, hard prompts35% weight, measured as gpt-6-luna-max, Sep 25, 202681
Search
basic search, plan exa-basic-payg
$70.00
Billable quantity
10,000 units
Fixed fee
$0.00
Usage
$70.00
Formula
fee + metered excess
Scraping
basic page, plan scraperapi-hobby-monthly
$49.00
Billable quantity
15,000 units
Fixed fee
$49.00
Usage
$0.00
Formula
fixed fee within hard cap

Fit scores are leaderboard percentiles: 90 means better than 90% of the models listed. Quality data: LMArena and Epoch AI, CC BY 4.0. Exact engine totals are summed before display rounding. Taxes, infrastructure and retry rates are excluded. Enabled affiliate links are marked individually and do not affect rankings. Disclosure

How the budget changes with usage

The same per-person workload at three audience sizes. Each row is recalculated, so providers and plans can change with volume.

Monthly users Model calls Searches / pages Monthly API estimate Try these inputs
1001,5001,000 / 1,500$26.68Calculate for 100 users →
1,00015,00010,000 / 15,000$125.75Calculate for 1,000 users →
10,000150,000100,000 / 150,000$1,511.50Calculate for 10,000 users →

Catalog 2026-09-26-7e5c5fac. Calculated when this page loads; an archived scenario may differ from current prices.

Define one research task before pricing it

For this example, one task is a bounded workflow: run the configured searches, read the configured pages, and use the configured model calls to produce an answer. The assumptions above define that workload. They are a planning example, not a claim about how every research agent behaves.

The model calls include all language-model work represented by the example, including the final answer. No additional report-writing call is added behind the scenes. If your agent uses a separate synthesis step, raise the call count or enter an exact workload that represents the total usage.

Searches and pages are both counted per task. Pages are not multiplied by searches a second time. If each search leads to several page reads in your implementation, add those reads together to obtain the pages-per-task input.

Three parts of the API budget

AI responses: the model processes context and generates text. Each call in this example uses a fixed input and output size. Scraped text and tool descriptions that you send to the model belong in the input estimate. A fixed-size model cannot reproduce every step of an agent with an expanding context window.

Web search: queries are priced using a supported search offer. A search result and a page retrieval are different operations, even when one leads to the other. Review the offer’s exact endpoint and billing assumptions. Tavily and Exa are examples in the catalog, with their sources and coverage stated on their provider pages.

Reading web pages: retrieval has its own billable quantity and plan. The example keeps page volume separate from model input tokens. Read the assumptions for Firecrawl and ScraperAPI before treating their included operations as interchangeable for your pages.

Why the next task may have a different effective cost

Model usage often scales with call and token volume, but fixed fees, included units, package increments and hard caps can change the effective cost of a task. The engine evaluates compatible plans at each usage level rather than multiplying one small example into a large forecast.

The breakdown above shows the chosen offers, the calculation formulas and the oldest relevant verification dates. A provider can become cheaper at a different volume. That does not establish that its search results or extraction quality fit your task; those qualities are not scored here.

The lowest-cost selection only covers compatible offers available in this catalog with fresh prices. An unavailable estimate is shown explicitly. It is never represented as a zero-cost research pipeline.

Adjust the workload in a useful order

Start with the number of people using the research feature and the tasks each person performs per month. These are monthly counts, not daily active users. Next, inspect searches and total pages per task. Finally, estimate how many model calls process those pages and how much text each call receives and produces.

Change one assumption at a time if you want to understand its effect. A broader search, a larger page sample and a longer report are different decisions. Increasing all three together makes it harder to tell what moved the budget.

Open the example in the calculator to try these changes. Keep a provider locked if you want to isolate volume changes from provider switching; unlock it when you want the engine to compare alternatives again.

A bounded estimate, not an autonomous-agent simulator

The example does not predict arbitrary tool loops, failed requests or retries. It does not automatically charge for embeddings, vector storage, browser compute, orchestration, hosting, taxes or engineering work. Cached-input and batch discounts are not applied. Recurring free allowances are off.

If real research runs vary greatly in depth, build separate light and heavy scenarios instead of relying on one average. A preliminary estimate cannot tell you whether the answer will be useful. Run a representative sample through the actual workflow, measure the billable operations, and update your inputs.

The methodology explains the supported billing rules and exclusions. If your app also answers ordinary support questions, the support-chatbot example covers that workload separately. Combine them only when the counts represent distinct operations.