Industry · 2026-06-10 · 8 min read
Goldman Sachs: AI Agents Will 24x Token Usage to 120 Quadrillion/Month by 2030 — and Reset Hyperscaler Cash Flow
Goldman Sachs Research forecasts agentic AI will multiply monthly token consumption 24x to 120 quadrillion tokens by 2030, while inference costs fall 60-70% per year — setting up a gross-margin inflection for NVIDIA, Microsoft, Google, AWS, Meta, OpenAI, and Anthropic. Here's what the forecast means for your token budget and how to plan around it.
TL;DR
- 24x token explosion: Goldman Sachs Research (analyst Jim Schneider) projects monthly token consumption rises from ~5 quadrillion today to 120 quadrillion tokens/month by 2030, driven primarily by agentic AI.
- Falling unit cost: Semiconductor providers are delivering 60-70% lower cost per inference token per year thanks to chip and data-center architectural gains.
- Margin inflection in 3-12 months: Hyperscaler gross margins are forecast to inflect upward as token revenue grows faster than capex amortization.
- Chip shortage: Expect a 12-18 month GPU/accelerator shortage before new fabs catch up.
- Enterprise lags consumer: Only 12% of knowledge workers will use agentic AI by 2030, rising to 37% by 2040 — the enterprise tail is long.
- What it means for you: Per-token prices will keep falling, but agent loops can blow up token counts 10-50x per task. Budget on agent-loop totals, not single-call prices.
The Headline Number: 120,000,000,000,000,000 Tokens/Month
Goldman Sachs Research modeled common consumer (online travel, smartphone takeover) and enterprise (customer service, coding) agentic use cases, then counted the tokens. The result: token consumption multiplies 24x between 2026 and 2030, with consumer agent workloads alone rising 12x.
For perspective, a single OpenAI ChatGPT query in 2024 averaged a few thousand tokens. An agent loop that books a flight chains dozens of LLM calls, plans, tool invocations, and reflections — Schneider notes agentic AI is "like taking a simple chatbot request and blowing it up 10-fold, 20-fold, 50-fold."
That math is exactly why we built the AI Agent Loop Cost Estimator — single-token prices look cheap until you multiply them by the real number of calls a working agent makes.
Why Cash Flow Inflects: The Two Curves Crossing
The Goldman thesis hinges on two curves moving in opposite directions:
Today, hyperscaler free cash flow is squeezed because capex on chips and data centers is eating revenue. Schneider's argument: as per-token revenue exceeds per-token cost, gross margins expand, operating cash flow grows, and the capex story shifts from "unsustainable burn" to "investing into operating leverage."
Semiconductor makers like NVIDIA are already running 70%+ gross margins. The inflection is at the *layer above* — the hyperscalers (Microsoft Azure, Google Cloud, AWS, Oracle) and model providers (OpenAI, Anthropic, Google DeepMind, xAI).
Who Wins, Who Waits
Not every AI workload benefits equally. Goldman flagged a real example where a real-time voice agent cost more in LLM tokens than a human agent because of latency requirements that force premium models on the critical path.
Highest leverage today:
- Coding agents — async, parallelizable, batch-friendly. (See: Hermes vs Paperclip vs OpenClaw cost comparison.)
- Text customer service — chatbot baseline is already efficient; agentic upgrades are incremental.
- Online travel & shopping — discrete tasks, tolerant of a few seconds latency.
Slower payoff:
- Real-time voice agents (latency penalty)
- Highly regulated enterprise workflows (compliance + integration tax)
- Small to medium businesses lacking internal AI ops capacity
Chip Shortage: The 12-18 Month Squeeze
Building a new fab takes ~3 years. Agent workloads emerged ~12 months ago. The industry is provisioning capacity sized for the 2025 chatbot world, not the 2026 agent world. Expect token prices on the newest accelerators (Blackwell, MI400, TPU v6) to stay firm even as legacy-chip prices crater. We've already tracked this divergence in our GPU pricing data.
What This Means for Your Token Bill
Three concrete planning shifts to make today:
1. Stop budgeting on per-call prices. Budget on agent-loop totals. A $3/1M input model called 40 times per task at 8K context = $0.96/task — not $0.024. Use the agent calculator before signing any contract.
2. Lean into the price decline. If inference cost is falling 60-70% per year, don't lock in 24-month committed-use discounts unless the discount exceeds your projected price decline. Most don't.
3. Stack caching + batching + routing. The compounding savings (covered in Reduce LLM Costs 50%) get *more* valuable as token volume scales, not less — even when unit prices fall, your bill is rising because volume is rising faster.
Track the Forecast on tokenscost.com
We update pricing for OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, Mistral, and others weekly, and the LLM Leaderboard shows daily token-usage shifts across providers. If Goldman's 24x forecast plays out, you'll see it in the leaderboard first — Schneider explicitly calls out that "every player will get dragged along to the upside, but at different rates." Different rates = visible divergence in real usage data.
Source
Goldman Sachs Research, "AI agents are forecast to boost tech cash flow as token usage soars" — interview with Jim Schneider, senior equity analyst, US semiconductor and IT services.