An AI agent can use more tokens and still cost less per completed task. It can also produce an impressively cheap answer that takes your team ten minutes to repair. Count the repair.
OpenRouter's agent usage figures make a useful starting point for an AI automation budget. They do not tell you what your next workflow will cost. For that, follow one job from the first model request to the person who accepts its output.
What the OpenRouter numbers actually show
In his analysis of OpenRouter traffic, Peter Walker reported roughly fourteenfold growth in agentic token usage from February 2026, compared with 2.8 times for human usage over the same period. Nearly 70% of agent token volume came from cached prompts.
Those are platform observations, not a census of business AI adoption. The classification uses signals such as tool calls, response intervals and turn counts. More agent tokens does not establish more dollars spent. Different models, cached input and generated output can all have different prices.
The practical implication is narrower and more useful. A growing agent workload needs a cost model that separates these categories. One total token counter will hide the decisions you need to make.
Follow a task instead of a chat message
Consider a support workflow that reads a ticket, looks up the customer's account, checks a policy and drafts a reply. A failed lookup may trigger another request. An unclear policy may send the case to a reviewer. Both still belong to the same job.

Give that job an identifier. Record every associated model call, external tool charge, retry and review. Then record the outcome. A draft is not a resolved ticket, and a successful API response is not proof that the answer was useful.
For workflows where steps are already predictable, compare the agent with conventional automation. A fixed rule does not need to think about whether to run itself.
Work through an AI agent cost example
Here is an illustrative monthly budget for 1,000 attempted support tasks. These rates and volumes are invented for the calculation, not a provider quote or a customer result. Assume no separately charged cache writes in this example.
| Cost item | Assumption | Monthly cost |
|---|---|---|
| Fresh input | 10 million tokens at $2 per million | $20 |
| Cached input | 30 million tokens at $0.20 per million | $6 |
| Generated output | 2 million tokens at $8 per million | $16 |
| External tools | 1,000 lookups at $0.01 each | $10 |
| Human review | 100 cases, 3 minutes each, $30 per hour | $150 |
| Total variable cost | Model, tools and review | $202 |
Suppose 900 tasks meet the agreed acceptance criteria. Variable cost per accepted task is $202 divided by 900, or about $0.224. Looking only at the $42 model bill would give about $0.047 per accepted task. That smaller figure is accurate for inference and misleading for the workflow.
Implementation, hosting, monitoring and ongoing maintenance are still outside this example. Add them separately. Also count failed attempts in the numerator, even when they never produce an accepted result.
Now change one assumption. If 300 cases need the same three-minute review, review costs $450 and the total becomes $502. With acceptance unchanged at 900 tasks, the unit cost rises to about $0.558. That is why a quality improvement can matter more than a cheaper token rate.
Cache stable context without expecting miracles
OpenRouter's prompt caching documentation distinguishes model and provider behaviour. Some caching is automatic. Some requires explicit configuration. Cache writes and reads can carry different charges.
Repeatedly sending a stable prompt prefix can allow cache reuse. Sending context again does not automatically forfeit the discount. Changing that prefix, switching providers or missing the retention window can affect reuse. Verify the behaviour of the actual route you use.
A high cache share also does not mean spend must grow more slowly than token volume. Double an unchanged mix of tokens at unchanged rates and the model bill doubles. Investigate unexpected cost per successful task, not two lines rising together.
Measure the calls and the outcome
OpenRouter's usage accounting documentation describes response usage fields for token counts, cost and caching details. Log the returned charge rather than estimating everything from a single advertised token rate. Keep model and provider identifiers with each request so a routing change is visible.
Your application must add the business context. Start with these five measures.
- Accepted tasks. Define what the workflow must achieve and who can reject its output.
- Cost per accepted task. Include unsuccessful attempts, tool calls and review time.
- Retry rate. Separate an expected retry from a loop that repeatedly fails for the same reason.
- Review burden. Measure actual minutes, not an optimistic percentage in a proposal.
- Completion time. Track slow cases as well as the typical run.
Set limits on calls, elapsed time and spend. Decide what happens when a limit is reached. Pausing with a useful explanation is usually preferable to paying an agent to rediscover that a database is unavailable.
Choose the next improvement from the evidence
If the model bill dominates, test shorter inputs, caching and a less expensive model on the same acceptance set. If review dominates, examine missing data, unclear instructions and recurring mistakes. If failures dominate, fix the integration before shopping for another model.
Use the AI Agent ROI Calculator to compare your own assumptions, then replace them with observed results from a small rollout. Our guide to moving enterprise AI from pilot to production covers the operational decisions around that rollout.
The budget worth defending is the one attached to useful work. A very busy agent is otherwise just another colleague with an impressive activity report.
