Three years ago inference was a third of AI compute. Today it's two-thirds. By 2029 it's three-quarters. Every analyst with a model — Deloitte, IDC, Gartner, McKinsey — agrees on the direction. They disagree only on the speed.
A clean trajectory, from a third to two-thirds in three years, and still climbing.
| Year | Inference share of AI compute | Training share | Source |
|---|---|---|---|
| 2023 | ~33% | ~67% | Deloitte TMT Predictions |
| 2024 | ~40% | ~60% | Industry estimate |
| 2025 | ~50% | ~50% | Deloitte TMT Predictions |
| 2026 | ~66% | ~34% | Deloitte / Gartner 55% IaaS |
| 2027 | ~70% | ~30% | Futurum, agent curve overlay |
| 2029 | ~65–75% | ~25–35% | Gartner IaaS forecast |
| 2030 | > 50% of all AI compute (and ~30–40% of total DC demand) | Steady ~30% | McKinsey |
The same conclusion, from three analysts with very different methodologies. That's the strongest possible signal in this market.
"Inference workloads will indeed be the hot new thing in 2026, accounting for roughly two-thirds of all compute (up from a third in 2023 and half in 2025). The market for inference-optimized chips will grow to over $50B in 2026."
55% of AI-optimized IaaS spending supports inference workloads in 2026, rising to over 65% by 2029. The market is reallocating cloud spend to serving, not training.
By 2029: 1B+ deployed agents, 217B+ actions/day, 3.7 TeraTokens/day consumed. Aggregate inference demand grows 1,000× by 2027 vs 2024 baseline. Per-token cost falls 87% — and total spend still triples.
The term now covers a much wider surface area than it did three years ago. Modern inference workloads include:
The original use case. Tight latency, streaming tokens, human-in-the-loop. Tens of milliseconds to first token, hundreds of tokens per second during decode.
Models that think longer to answer better — o1, o3, Claude with extended thinking, DeepSeek R1. Each request may consume 10–100× more compute than a standard chat call.
Multi-step, tool-using, long-context. Agents call models dozens of times per task, each call potentially long-context. Bursty, latency-sensitive, cache-unfriendly.
Vector embeddings, reranking, hybrid search. Lower per-call compute, but massive volume — every RAG pipeline, every semantic search, every recommendation system.
Vision, audio, video understanding and generation. Sora, Veo, image models, voice agents. Different shape of workload, different rack profile.
Code, biotech, materials, legal, finance. Specialized models deployed for specific enterprise workflows. Often on-prem, often latency-bound.
"Inference already accounts for roughly 85% of enterprise AI spend in 2026, up from around half of all AI compute just two years ago. Inference demand will outpace training by 118× by the end of 2026."
— IDC FutureScape 2026 synthesis (agentmarketcap.ai)"By 2030, AI inference workloads make up more than 40 percent of total data center demand, with non-AI workloads dropping below one-third, and AI training holding steady at just under 30 percent."
— McKinsey, The Future of AI Workloads