Why Trimming Tokens Can Raise Your Agent Bill

Agents don’t buy tokens, they rent memory. Part 1 of 2 on agent economics: why the cost gap between agents is architectural.

Researchers cut a coding agent’s tool output by 38%. Its bill went up by 6.8%.

This is not an anecdote. It comes from a July 2026 study of nearly 2,900 provider-billed Claude Code runs across 103 tasks and three Claude models, with each compression setup paired against a baseline on the same task (Weinberger & Hozez, arXiv 2607.12161). The setup that removed about 38% of raw tool-output tokens raised billed cost by 6.8%, with a 95% confidence interval of +2.8% to +11.3%. On one benchmark subset, task completion fell as well. The authors work at PointFive, which sells AI and cloud cost-management tools, and they note the limits themselves: one agent harness, one provider’s prices and cache rules, one point in time.

If you manage an agent budget, that result should bother you. Nearly every cost playbook in circulation says the same thing: fewer tokens, smaller bill. For agents, that is wrong often enough to matter. And the reason connects to the number hiding under this summer’s most-shared AI chart.

The headline everyone quoted, and the number under it

In August, a16z’s Charts of the Week circulated OpenRouter data showing agents now use nearly five times the tokens humans do: about 7.3 trillion tokens (seven-day average) for agentic traffic against about 1.4 trillion for human traffic, with agentic volume up roughly 14x since February (a16z). OpenRouter’s own analysis puts a single agentic request at about 15x the tokens of a human one (OpenRouter).

Four caveats before using that number. OpenRouter is one gateway, not the market. It labels traffic as agentic or human per API key, using a weighted behavioural score rather than ground truth. Token volume is not spend. And a16z is an OpenRouter investor.

The more interesting figure sits underneath. Most agent tokens are cache reads: inputs the provider has already processed and serves at a discount. a16z puts it at more than 85% of agentic token volume, and notes that cached tokens account for nearly all of the growth. In plain terms: agents spend most of their tokens re-reading what they have already seen.

It is tempting to read this as a quirk of how agents behave. It is arithmetic.

An agent loop re-reads itself, so cost grows with the square of its length

Every turn of an agent loop sends the model the full context so far: system prompt, tool definitions, prior steps, tool results. The model reasons, calls a tool, and the result is appended. On the next turn, everything goes back in again.

Run the numbers on a modest agent:

  • a 20,000-token system prompt and tool definitions
  • 2,000 new tokens added per turn (a tool call and its result)
  • 50 turns to finish the task

The agent is billed for 3.45 million input tokens. Only about 118,000 of them are new. Roughly 96.6% of the input is re-reading.

The shape matters more than the totals. Total input grows with the square of the number of turns, because each turn re-reads everything before it. Double the task to 100 turns and the agent is billed for 11.9 million input tokens: 3.45x the input for 2x the work. Even with caching, which flattens the price, the task in this example costs about 2.8x as much.

This is the shape of a simple loop that keeps its whole history. Agents that compact their context or hand work to sub-agents flatten the curve, and that is exactly why those techniques exist. But the default loop is quadratic, so the 85%-plus cached share is not a sign of unusual agent behaviour. It is what a multi-turn loop produces. Real agents land below the 96% in this example because of compaction, sub-agents and short tasks, and in the same region because the underlying loop has the same shape.

Infographic: over 50 turns an agent is billed for 3.45 million input tokens while only 118,000 are new, so 96.6% of input is re-reading. Caching cuts the price but not the curve; cutting tool output 38% raised the bill 6.8%; the cost gap lives mainly in loop design and cache discipline.
The worked example, the price effect of caching, the 38% paradox, and where the cost gap lives.

Caching bends the price, not the curve

Prompt caching is what keeps this affordable. If the start of the context is byte-for-byte identical to what the provider saw last turn, the provider reuses its stored computation (the KV cache) and charges a fraction of the normal input price.

The discount is large. Manus, which runs a production agent at an input-to-output ratio of about 100:1, reported cached input on Claude Sonnet at $0.30 per million tokens against $3.00 uncached, a 10x gap. It called KV-cache hit rate “the single most important metric for a production-stage AI agent” (Manus). Current price cards keep the same ratio. In August 2026, gpt-5.6-luna listed $0.20 per million uncached input tokens, $0.25 to write the cache, $0.02 to read it, and $1.20 for output (arXiv 2609.19607).

Apply those rates to the 50-turn example, including the premium for writing new tokens to the cache, and the input cost drops about 7x. That is a big saving. But it is a lower price on the same curve, not a different curve. The agent’s cost still grows with the square of its length, just from a smaller base.

Caching also changes where the money goes, in three ways:

  • Output and reasoning tokens become the dominant line. At $1.20 against $0.02, an output token costs 60 times a cached read. Once caching works, the tokens the model writes, including invisible reasoning tokens billed as output, often outweigh everything it reads.
  • Writing the cache costs extra. On some providers the first write is priced above normal input, so a cache that is written and never re-read is a loss.
  • Caches expire. Cache lifetimes are typically measured in minutes. If an agent pauses longer than that, say while waiting for human approval, the next turn re-pays full price for the entire accumulated context. That is a cost cliff, and it comes from workflow design, not token count.

One clarification, because the terms get blurred: prompt (prefix) caching is the provider reusing computation for an identical prefix, priced per read. Explicit context caching is a named cache object you create and pay to store over time. Semantic or response caching happens in your own application and returns a stored answer without calling the model at all. They have different economics, and a budget should treat them separately.

So compression attacked the wrong cost line

Back to the paradox. The Claude Code study’s own cost breakdown explains it. Prompt-cache traffic made up most of the input-side spend. Tool-output compression could only touch the fresh tokens entering each turn, which is a small slice of the bill once caching is working.

Then the second-order effect arrived. With less detail in its tool results, the agent changed its behaviour: it did more retrieval, more diagnosis, more testing, and took more turns. Each extra turn re-read the entire accumulated context. A small saving on fresh tokens was overwhelmed by a larger cost from a longer loop.

The authors conclude that token count is an inadequate proxy for cost in tool-heavy agents, and that efficiency has to be measured as success-adjusted, end-to-end cost. The general lesson for anyone running agents: any change that lengthens the loop, or breaks the cached prefix, can cost more than it saves. Trimming tokens is not neutral. It is a change to the agent’s trajectory, and the trajectory is where the cost is.

Routing has the same blind spot

Model routing, sending easy requests to cheap models and hard ones to expensive models, is the other standard cost lever. The peer-reviewed results are real but narrower than the marketing. RouteLLM, the most-cited study, reported cutting cost by up to 85% against an all-GPT-4 baseline while keeping 95% of its quality, but that was on a chat benchmark. On maths and knowledge benchmarks its cost-saving ratios were about 1.4x to 1.5x, at 87–92% of GPT-4’s quality (arXiv 2406.18665). The gain depends heavily on how many requests are genuinely easy.

Those results were measured on single-turn chat. In an agent, routing collides with caching. A cache belongs to one model. Switch models mid-session and the new model has no cache for the accumulated context, so the next turn pays full price on everything. A router that saves 40% on per-token price can lose it to one forced cache miss on a 100,000-token context.

The design rule follows directly: route per sub-agent or per session, never per turn. Give a cheap model its own short-lived context for a bounded sub-task, and keep the long-running loop on one model.

The 10x gap lives in the loop design

If token trimming is unreliable and routing is bounded, where do the large differences between agents come from? The evaluation research points one way: architecture.

The clearest recent evidence comes from Princeton’s Holistic Agent Leaderboard, which ran 21,730 agent rollouts across 9 models and 9 benchmarks (HAL, arXiv 2510.11977). On one web-navigation benchmark, two agent setups, each a different framework paired with a different model, landed 9x apart in cost ($171 against $1,577) for a two-point difference in accuracy. The authors’ conclusion is that the choice of framework is consequential for both cost and accuracy. The same study found the most expensive model sat on the cost-accuracy frontier in only 1 of 9 benchmarks. And in 21 of 36 model-framework-benchmark combinations, raising the reasoning effort produced equal or lower accuracy. More thinking is not reliably better. It is reliably more expensive.

The pattern predates this year’s models. In 2024, the same Princeton group showed that a simple retry-with-rising-temperature baseline on HumanEval scored 93.2% at $2.45, while the elaborate LATS tree-search agent scored 88.0% at $134.50. Higher accuracy at about one-fiftieth of the cost (Kapoor et al., arXiv 2407.01502). A separate enterprise study reports accuracy-maximising configurations costing 4.4–10.8x more than cost-aware alternatives with comparable performance (CLEAR, arXiv 2511.14136). It is a single-author preprint, but it points the same way.

Put the evidence side by side and a rough ranking emerges. It is a synthesis, not a measurement: the figures come from different studies, tasks and years, so treat the order as more reliable than the multiples.

  1. Loop design and framework: roughly 4x to 50x, from how many turns, what structure, what retries.
  2. Cache discipline: up to about 10x on input, from whether the prefix stays stable.
  3. Routing: large on easy chat traffic, modest on harder work, and smaller still if it breaks the cache.
  4. Token compression: small, and measured as negative in at least one careful study.

Levers three and four get most of the attention, because they can be bolted on after the fact. The money is in one and two, which have to be designed in from the start.

What to instrument on Monday

If the cost gap is architectural, cost control starts in the design review, not the finance review. Five metrics belong on every production agent’s dashboard, next to accuracy:

  • Cost per successful task, not cost per token or per request. An agent that succeeds 60% of the time effectively costs 1.67x its per-run price.
  • Cache-hit rate, as a percentage of input tokens. A sudden drop means someone broke the prefix.
  • Turns per task, because the cost grows with the square of it.
  • Output and reasoning share of cost, the line that grows once caching works.
  • 95th-percentile cost per task, because the expensive tail is where budgets break.

And four design rules that are almost free if adopted early and expensive to retrofit:

  • Keep the prompt prefix stable. No timestamps, session IDs or per-request variables at the top of the context. One changed token invalidates the cache from that point on.
  • Make context append-only. Do not rewrite earlier steps. Summarise only at deliberate checkpoints, knowing compaction resets the cache.
  • Serialise deterministically. Tool definitions and JSON with unstable key order silently break caching.
  • Isolate, don’t switch. Use sub-agents with their own contexts rather than changing models inside a long-running loop.

A caveat: not every agent deserves this attention. If an internal agent spends a few hundred dollars a month, an engineer-week costs more than any saving. A reasonable test is whether monthly spend, times a realistic 30–50% reduction, times twelve months exceeds the engineering cost. Below that line, adopt the four design rules, which cost almost nothing, and move on.

The bill is decided before finance sees it

The lesson of the 38% that raised the bill is not that compression is bad. It is that token count is the wrong unit for agent economics. An agent’s cost is set by the shape of its loop and the stability of its context, decisions made by architects long before anyone opens an invoice.

Engineers can bend the curve. But a well-designed agent still has an uneven, heavy-tailed cost profile, and efficiency gains have a habit of turning into more usage rather than lower bills. Deciding how much to spend, how to cap the tail, and what a successful task is worth is a different discipline. That is Part 2: Budgeting for Agents: Spend Envelopes, Tail Risk, and Cost per Outcome.


Reference: the twelve levers of agent cost

#LeverWhat it controlsWatch for
1Model tier (frontier / mid / open-weight / self-hosted)Price per tokenCompare per successful task; a cheap model that needs more turns can cost more
2Routing and cascadesBlended priceModel switches destroy the cache; route per sub-agent or session
3Prompt-caching disciplineCache-read shareStable prefix, append-only context, deterministic serialisation
4Cache lifetime and write premiumWrite cost vs read savingsPauses longer than the cache lifetime re-price the whole context
5Context management (compaction, sub-agents, retrieval)Total input volumeCompaction resets the curve but also the cache; time it deliberately
6Output and reasoning budgetThe largest line once caching worksMore reasoning often buys no accuracy
7Loop shape (turns, retries, fan-out)The multiplier on everythingHard iteration caps are circuit breakers
8Success rate and reworkDivides the whole costHuman review time is usually unmodelled
9Pricing mode (on-demand / batch / committed / flat-rate plans)Price per unit of workBatch for latency-tolerant agents; flat-rate vs metered arbitrage
10Self-hostingCapex and opex vs API spendFor agents, KV-cache memory, not compute, is the binding constraint
11Non-token costs (search, sandboxes, vector stores, observability)InfrastructureScales with tool calls, not tokens
12Variance and governance (caps, chargeback, alerts)Tail riskBudget the 95th percentile, not the mean

Sources

  • Weinberger & Hozez, Token Reduction Is Not Cost Reduction, arXiv 2607.12161 (July 2026)
  • OpenRouter, DeepSeek V4 Is Earning Agentic Token Share (June 2026)
  • a16z, Charts of the Week: Winds of Thematic Change (21 Aug 2026), OpenRouter data, charts by Peter Walker
  • Manus, Context Engineering for AI Agents: Lessons from Building Manus
  • DeltaSelect, arXiv 2609.19607 (gpt-5.6-luna rates as of 16 Aug 2026; same rates listed after OpenAI’s 30 July price change)
  • Ong et al., RouteLLM, arXiv 2406.18665
  • Kapoor et al., Holistic Agent Leaderboard, arXiv 2510.11977
  • Kapoor et al., AI Agents That Matter, arXiv 2407.01502
  • Mehta, Beyond Accuracy (CLEAR), arXiv 2511.14136