For this example, one task is a bounded workflow: run the configured searches, read the configured pages, and use the configured model calls to produce an answer. The assumptions above define that workload. They are a planning example, not a claim about how every research agent behaves.
The model calls include all language-model work represented by the example, including the final answer. No additional report-writing call is added behind the scenes. If your agent uses a separate synthesis step, raise the call count or enter an exact workload that represents the total usage.
Searches and pages are both counted per task. Pages are not multiplied by searches a second time. If each search leads to several page reads in your implementation, add those reads together to obtain the pages-per-task input.
AI responses: the model processes context and generates text. Each call in this example uses a fixed input and output size. Scraped text and tool descriptions that you send to the model belong in the input estimate. A fixed-size model cannot reproduce every step of an agent with an expanding context window.
Web search: queries are priced using a supported search offer. A search result and a page retrieval are different operations, even when one leads to the other. Review the offer’s exact endpoint and billing assumptions. Tavily and Exa are examples in the catalog, with their sources and coverage stated on their provider pages.
Reading web pages: retrieval has its own billable quantity and plan. The example keeps page volume separate from model input tokens. Read the assumptions for Firecrawl and ScraperAPI before treating their included operations as interchangeable for your pages.
Model usage often scales with call and token volume, but fixed fees, included units, package increments and hard caps can change the effective cost of a task. The engine evaluates compatible plans at each usage level rather than multiplying one small example into a large forecast.
The breakdown above shows the chosen offers, the calculation formulas and the oldest relevant verification dates. A provider can become cheaper at a different volume. That does not establish that its search results or extraction quality fit your task; those qualities are not scored here.
The lowest-cost selection only covers compatible offers available in this catalog with fresh prices. An unavailable estimate is shown explicitly. It is never represented as a zero-cost research pipeline.
Start with the number of people using the research feature and the tasks each person performs per month. These are monthly counts, not daily active users. Next, inspect searches and total pages per task. Finally, estimate how many model calls process those pages and how much text each call receives and produces.
Change one assumption at a time if you want to understand its effect. A broader search, a larger page sample and a longer report are different decisions. Increasing all three together makes it harder to tell what moved the budget.
Open the example in the calculator to try these changes. Keep a provider locked if you want to isolate volume changes from provider switching; unlock it when you want the engine to compare alternatives again.
The example does not predict arbitrary tool loops, failed requests or retries. It does not automatically charge for embeddings, vector storage, browser compute, orchestration, hosting, taxes or engineering work. Cached-input and batch discounts are not applied. Recurring free allowances are off.
If real research runs vary greatly in depth, build separate light and heavy scenarios instead of relying on one average. A preliminary estimate cannot tell you whether the answer will be useful. Run a representative sample through the actual workflow, measure the billable operations, and update your inputs.
The methodology explains the supported billing rules and exclusions. If your app also answers ordinary support questions, the support-chatbot example covers that workload separately. Combine them only when the counts represent distinct operations.