Analysis · 2026-10-04 · 8 min read
AI Agents Now Use 5x More Tokens Than Humans — What That Means for Your Bill
OpenRouter data shows agents burning 7.3 trillion tokens a week versus 1.4 trillion for humans, with 85% of agent traffic coming from cached prompts. Here is the TokenCost breakdown of what that means for API budgets.
The Machines Are the Customers Now
Futurum Group CEO Daniel Newman put it bluntly in a recent X post: "AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher." The data behind the claim is an Andreessen Horowitz (a16z) chart built on OpenRouter figures, first reported by Tom's Hardware: as of August 2026, agents routed through OpenRouter consumed 7.3 trillion tokens on a 7-day average, versus 1.4 trillion for humans.
The crossover happened in February 2026 — just six months before agent traffic hit 5x human volume. Since that crossover, agentic usage is up 14x while human usage grew only 2.8x.
The Number That Matters More: 85% Cached
Here is the detail most headlines skip: more than 85% of agent tokens come from cached prompts. Agents are mostly rereading what they have already seen — reloading system prompts, tool definitions, memory of prior steps, and schema constraints on every single turn of their loop.
That is not waste in the naive sense; it is how agent architectures work. But it has two enormous implications for cost:
1. Cached-input pricing is now the most important line on the price sheet. A model with a cheap headline input price but no cache discount is far more expensive for agent workloads than a model with 80-90% cached-input rates.
2. Token volume is decoupling from value. If agents 10x their token burn by rereading context, your bill does not have to 10x — but only if you are capturing cache discounts.
The TokenCost Math
Run the numbers on a realistic agent workload: 1 billion input tokens per month, 85% cacheable, 100 million output tokens.
The spread is 10x between a frontier model and a well-cached mid-tier model on identical agent traffic. At the 5x-agents scale Newman describes, model routing stops being an optimization and becomes the difference between a viable product and an unsustainable one.
Why Agents Reread: The Structural Reasons
- Stateless APIs. Every agent turn resends the full conversation, tool schemas and system prompt. A 50-turn agent session resends its instructions 50 times.
- Tool-heavy loops. Browser snapshots, file trees and API responses are chunky payloads that get folded into context and re-read on subsequent turns.
- Delegation fan-out. Orchestrators spawn sub-agents, each with its own copy of the shared context.
OpenRouter classifies each API key as agentic, mixed, or human using a 7-signal weighted score (tool call rate, turn count, gap timing and more), so the 5x figure is measured behavior, not a survey estimate.
The Hardware Angle: KV Cache and HBM
The Tom's Hardware report flags a second-order effect: all that cached context has to live somewhere. Models hold stored context in the KV cache, and a16z ties the cached-prompt explosion directly to rising demand for high-bandwidth memory (HBM) — worsening an already tight RAM market. NVIDIA's data center revenue surged 117% year over year to $89 billion in its latest quarter, and agentic inference is a growing share of that demand. Expect memory pressure to keep cached-input discounts generous (providers want cache hits) while long-context premium tiers stay expensive.
Five Things to Do About It
1. Audit your cache hit rate. If your agent workload is below 70% cached input, you are leaving the biggest discount on the table. Keep system prompts and tool schemas byte-stable.
2. Route by task, not by habit. Use frontier models only for planning and verification steps; run extraction, classification and tool-use turns on tier-1 or tier-2 models. See our three-tier market breakdown.
3. Cap agent loops. Set hard turn limits and token budgets per session. An unbounded agent loop at 5x human volume is an unbounded bill.
4. Trim context before price cliffs. Several models (including Grok 4.7 at 200K tokens) double their rates past a context threshold. Summarize or retrieve instead of accumulating.
5. Track cost per completed task, not per token. Agents retry. The cheapest model per token is often not the cheapest per successful outcome.
Bottom Line
The 5x figure will look quaint within a year — Newman expects 10x, and the February-to-August growth curve supports him. The winners in this market are not the teams using the most tokens; they are the teams whose agents reread context at a 90% discount on a $0.30/M model instead of full price on a $10/M one. Check current cache rates on our live pricing table and compare value scores on the LLM leaderboard.
Sources
- Tom's Hardware — AI agents use 5x more tokens than humans
- a16z Charts of the Week — OpenRouter token usage data
- OpenRouter — model gateway and routing platform
Related Reading
- Prompt Caching Explained — How to capture the 80-90% cached-input discount agents depend on.
- AI Token Pricing in 2026 — The three-tier market and a six-step routing policy.
- The Context Window Cost Trap — Why bigger context is not cheaper context.