The product category is no longer a forecast. NVIDIA's GB200 NVL72 and GB300 NVL72 ship today as inference racks first, training platforms second. 72 GPUs, 1.8 TB/s per-GPU NVLink, 30 TB of fast memory, 130 TB/s aggregate bandwidth — engineered for the agent economy.
Three product lines that have already converged on rack-as-inference-unit.
72 Blackwell GPUs, 36 Grace CPUs, 130 TB/s NVLink fabric, 30× faster real-time LLM inference vs H100, 1.8 TB/s per-GPU bandwidth. The first product explicitly positioned as a real-time inference rack at trillion-parameter scale.
The follow-on. 1.5× the AI performance of GB200. Purpose-built for "test-time scaling inference and AI reasoning tasks." Paves the way for what NVIDIA calls "the age of AI reasoning."
The smaller cousin: 11× faster LLM inference vs Hopper. For inference workloads that don't need 72-GPU NVLink domains, but still benefit from rack-scale engineering and shared memory.
A standard GB200 NVL72, by the numbers.
| Component | GB200 NVL72 specification |
|---|---|
| GPUs | 72× NVIDIA Blackwell GPUs |
| CPUs | 36× NVIDIA Grace (Arm Neoverse V2) |
| CPU cores | 2,592 Arm cores |
| GPU memory | 13.4 TB HBM3e, 576 TB/s aggregate bandwidth |
| Total fast memory | ~30 TB (HBM + LPDDR5X) |
| NVLink fabric | 130 TB/s aggregate, 1.8 TB/s per GPU |
| Domain size | Acts as a single, massive GPU (576 GPUs via NVLink switch) |
| Peak AI performance | 1,440 PFLOPS NVFP4 / 720 PFLOPS dense |
| Cooling | 100% liquid-cooled |
| Rack power | ~120 kW |
Hardware is half the story. The other half is the inference serving layer.
The inference orchestration platform. Disaggregated prefill and decode, autoscaling, KV-cache offload, GPU multiplexing. Designed specifically to make the NVL72 domain behave as a single inference engine.
NVIDIA's LLM inference engine. Optimized kernels for transformer decoding, in-flight batching, speculative decoding, LoRA serving. Pairs with Dynamo and the NVL72 fabric for MoE-scale inference.
The open-source inference stack. PagedAttention, RadixAttention, continuous batching — the techniques that made high-throughput open-source inference possible. They all run on inference racks, at scale.
Langfuse, Helicone, Arize, WhyLabs, plus cloud-native options. Per-token, per-request, per-model cost and quality tracking. Without this layer, an inference rack is a meter running at 1,000× speed.
The first wave of inference rack deployments, named publicly.
For a generation, the unit of AI design was the chip. H100, A100, MI300 — chip comparisons were the conversation. The rack changes that. Inference workloads need coherent memory, low-latency interconnect, and a serving layer that's tuned as a system, not a part. NVIDIA's most-quoted numbers are no longer "TOPS per GPU" but "30× real-time LLM inference" and "1,000× speedup at the rack level." The rack is the product.
And because inference is the dominant workload, and the rack is the unit of inference, the rack is the unit of the AI economy. That makes InferenceRacks.com the address of the market — not a metaphor for it.
"We designed Blackwell Ultra for this moment — it's a single versatile platform that can easily and efficiently do pretraining, post-training and reasoning AI inference."
— Jensen Huang, NVIDIA, GTC 2025 keynote"The combination of NVIDIA Dynamo and NVIDIA GB200 NVL72 creates a powerful compounding effect that optimizes inference performance for AI factories that are deploying MoE models, such as DeepSeek R1 and the newly released Llama 4 models in production."
— NVIDIA, on the inference rack + serving stack"Oracle's state-of-the-art GB200 deployment... delivers exceptional performance and energy efficiency for agentic AI powered by advanced AI reasoning models."
— NVIDIA & Oracle, on the first wave of Blackwell inference racks