Training a frontier model costs tens of millions of dollars. Serving it to a billion agents, continuously, 24/7, costs a hundred times more — every year, for the life of the model. That's why 80–90% of every AI system's lifetime compute spend is inference.
Same model, same vendor, real numbers. Training is the spike. Inference is the floor.
| Dimension | Training | Inference |
|---|---|---|
| When it runs | Once per model version — days or weeks | Every request, 24/7, indefinitely |
| Accounting treatment | One-time capital event / R&D | Recurring operating expense |
| Per-run cost | $10M – $500M+ for a frontier run | ¢ to $ per request, $M to $B per year |
| Share of lifetime compute | 10 – 20% | 80 – 90% |
| Cost trend | Climbing at the frontier (bigger models, more tokens) | Per-token cost falling 9× to 900× per year |
| Primary cost driver | Model size and dataset scale | Token volume × request latency |
| Workload pattern | Bursty, batch, predictable | Continuous, bursty, latency-bound, geo-distributed |
| Geography | One cluster, one site, one power deal | Many regions, edge, on-prem, near user |
| Buyer inside the enterprise | Head of AI / R&D | Head of Engineering, Product, CX, FinOps |
The economics of training and inference are different enough that the rack profile is different too.
Maximizes per-GPU FLOPs and interconnect bandwidth for large all-reduce operations. Runs one job, saturates the fabric, then idles. Optimized for batch throughput, not latency. Liquid-cooled, megawatt-class.
Maximizes tokens-per-dollar-per-watt at target latency. Runs millions of requests concurrently, dynamically batched. Optimized for KV-cache capacity, prefill/decode disaggregation, and per-request SLA. Often smaller, denser, more numerous.
Modern AI factories run both on the same rack platform — the NVL72 is explicitly sold as "training and inference in one." The fact that it can do both is a feature, not a product identity. The product identity is "the rack."
The product identity is the inference rack, because inference is 80–90% of the spend. Even a rack that does both will be marketed, named, and remembered for its inference profile. That's the brand. That's the URL.
A worked example: an enterprise that trains a $30M frontier model in 2026. How much does it spend serving that model over its lifetime?
One frontier training run, 2026. Single line item. Capitalized or expensed depending on accounting.
Initial rollout, mostly internal users, copilots, a handful of agent pilots. Typical 2024–2025 baseline.
Department-scale agents, customer-facing products, embedded AI. Typical 2026–2027 enterprise projection.
Enterprise-wide agent fleets, AI as the primary interface. Typical 2028–2029 projection at large enterprises.
The model gets retrained every 6–18 months at the frontier. So training spend reappears — but still as ~10–20% of total.
Over 4 years, the typical large-enterprise AI system spends ~10× more on inference than training. Some models will hit 100×.
"Inference now consumes the majority of an AI system's lifetime compute, with industry analyses putting it at roughly 80 to 90% of total compute dollars over a model's lifecycle versus 10 to 20% for training."
— Telnyx, AI training vs inference: the 2026 economics split"The AI inference market will grow from $106 billion in 2025 to $255 billion by 2030, with a 19.2% compound annual growth rate. Gartner projects that 55% of AI-optimized IaaS spending will support inference workloads in 2026, reaching over 65% by 2029."
— Introl / Gartner, AI Inference vs Training Infrastructure economics