Unit economics

Per-token cost falls 10× a year.
Total spend still grows 100×.

Inference is the only market in technology where the price is collapsing 300–360× in three years and the dollar volume is exploding at the same time. The deflation of intelligence is fueling the inflation of intelligence budgets.

The deflation curve

Per-million-token cost across the frontier, 2023 → 2026.

Model Launch Input price / 1M tokens Output price / 1M tokens
GPT-4 (initial) March 2023 $30.00 $60.00
GPT-4 Turbo Nov 2023 $10.00 $30.00
GPT-4o May 2024 $5.00 $15.00
GPT-4o (latest) 2025 $2.50 $10.00
Claude 3.5 Sonnet 2024 $3.00 $15.00
DeepSeek V3 / R1 Late 2024–2025 $0.14 – $0.55 $0.28 – $2.19
Gemini 2.5 Flash-Lite April 2026 (market floor) $0.10 $0.40

Sources: OpenAI, Anthropic, DeepSeek, Google pricing pages, May 2026. Prices are public list prices; enterprise contracts differ.

The budget paradox

Deflation × Volume = Inflation. The single most important equation in the AI economy right now.

−10×

Per-token cost per year

The market floor for capable model inference falls roughly 10× per year. In 37 months, GPT-4-class output went from $60/M tokens to $0.40/M tokens. 150× in three years.

+1,000×

Token volume by 2027

IDC: token and API-call loads grow 1,000× by 2027. Agents get hungrier as they get more capable, and there are 10× more of them.

+100×

Aggregate spend growth

10× cheaper × 1,000× more = 100× aggregate spend growth, even after the 87% per-action efficiency gain. The per-unit opex falls. The total opex explodes.

$68B

Annual token delivery, 2029

The 100×-spend curve terminates in a $68B+ annual market for token delivery alone. Add GPU rental, networking, observability, FinOps tools — the addressable market is much larger.

What enterprise inference budgets look like

Inside the Fortune 500, the line item is moving fast. Per-agent-platform estimates, all referenced against IDC's forecast.

Period Large-enterprise annual inference spend What it buys
2024–2025 (baseline) $5M – $15M Early copilots, internal chatbots, a handful of agent pilots
2026–2027 (projected) $50M – $150M Department-scale agents, customer-facing copilots, embedded AI in products
2028–2029 (projected) $500M – $1.5B Enterprise-wide agent fleets, autonomous workflows, AI as primary interface

Source: IDC FutureScape 2026, agent-economy forecast, scaled to large-enterprise baseline. Actual spend varies dramatically by industry, deployment model, and model mix.

The FinOps discipline emerges

When a line item grows 100× in 4 years, you can't manage it with spreadsheets. A whole product category — AI FinOps — has emerged to handle it.

📊

Per-token observability

Per-model, per-prompt, per-agent, per-workflow cost tracking. New vendors and new open-source tools (Langfuse, Helicone, OpenLLMetry) are building this layer in real time.

🔀

Model routing & cascading

Route cheap requests to cheap models, expensive ones to the frontier. GPT-4-class problems get GPT-4-class compute; everything else gets Gemini Flash-Lite. Same answer, 1/50th the cost.

Batching, caching, autoscaling

Continuous batching, prompt caching, prefix caching, speculative decoding. Each one shaves another 2–5× off the cost of a given workload. The rack now ships with these features baked in.

"AI agent inference cost is deflating 10× annually — faster than PC compute or dotcom bandwidth — but aggregate enterprise spend is increasing 3–5× because agent demand growth outpaces cost decline."

Perea.ai, Agent Inference Unit Economics research

"The math: 10× cheaper per token + 1,000× more tokens = 100× aggregate spend growth, even with 87% per-action efficiency gains."

IDC FutureScape 2026 synthesis

The unit economics are unambiguous.

Per-token cost falls. Token volume explodes 1,000×. The market grows 100×. The unit of that growth is the inference rack. InferenceRacks.com.

Acquire the domain →