The inference rack

The inference rack exists. It's called NVL72.

The product category is no longer a forecast. NVIDIA's GB200 NVL72 and GB300 NVL72 ship today as inference racks first, training platforms second. 72 GPUs, 1.8 TB/s per-GPU NVLink, 30 TB of fast memory, 130 TB/s aggregate bandwidth — engineered for the agent economy.

The inference rack, in production today

Three product lines that have already converged on rack-as-inference-unit.

72

GB200 NVL72

72 Blackwell GPUs, 36 Grace CPUs, 130 TB/s NVLink fabric, 30× faster real-time LLM inference vs H100, 1.8 TB/s per-GPU bandwidth. The first product explicitly positioned as a real-time inference rack at trillion-parameter scale.

72+

GB300 NVL72 (Blackwell Ultra)

The follow-on. 1.5× the AI performance of GB200. Purpose-built for "test-time scaling inference and AI reasoning tasks." Paves the way for what NVIDIA calls "the age of AI reasoning."

16

HGX B300 NVL16

The smaller cousin: 11× faster LLM inference vs Hopper. For inference workloads that don't need 72-GPU NVLink domains, but still benefit from rack-scale engineering and shared memory.

What an inference rack actually contains

A standard GB200 NVL72, by the numbers.

Component GB200 NVL72 specification
GPUs 72× NVIDIA Blackwell GPUs
CPUs 36× NVIDIA Grace (Arm Neoverse V2)
CPU cores 2,592 Arm cores
GPU memory 13.4 TB HBM3e, 576 TB/s aggregate bandwidth
Total fast memory ~30 TB (HBM + LPDDR5X)
NVLink fabric 130 TB/s aggregate, 1.8 TB/s per GPU
Domain size Acts as a single, massive GPU (576 GPUs via NVLink switch)
Peak AI performance 1,440 PFLOPS NVFP4 / 720 PFLOPS dense
Cooling 100% liquid-cooled
Rack power ~120 kW

The software stack that makes it an inference rack

Hardware is half the story. The other half is the inference serving layer.

NVIDIA Dynamo

The inference orchestration platform. Disaggregated prefill and decode, autoscaling, KV-cache offload, GPU multiplexing. Designed specifically to make the NVL72 domain behave as a single inference engine.

🧠

TensorRT-LLM

NVIDIA's LLM inference engine. Optimized kernels for transformer decoding, in-flight batching, speculative decoding, LoRA serving. Pairs with Dynamo and the NVL72 fabric for MoE-scale inference.

🌐

vLLM, SGLang, llama.cpp

The open-source inference stack. PagedAttention, RadixAttention, continuous batching — the techniques that made high-throughput open-source inference possible. They all run on inference racks, at scale.

🔍

Observability & FinOps

Langfuse, Helicone, Arize, WhyLabs, plus cloud-native options. Per-token, per-request, per-model cost and quality tracking. Without this layer, an inference rack is a meter running at 1,000× speed.

What hyperscalers are buying

The first wave of inference rack deployments, named publicly.

Why the rack, not the chip

For a generation, the unit of AI design was the chip. H100, A100, MI300 — chip comparisons were the conversation. The rack changes that. Inference workloads need coherent memory, low-latency interconnect, and a serving layer that's tuned as a system, not a part. NVIDIA's most-quoted numbers are no longer "TOPS per GPU" but "30× real-time LLM inference" and "1,000× speedup at the rack level." The rack is the product.

And because inference is the dominant workload, and the rack is the unit of inference, the rack is the unit of the AI economy. That makes InferenceRacks.com the address of the market — not a metaphor for it.

"We designed Blackwell Ultra for this moment — it's a single versatile platform that can easily and efficiently do pretraining, post-training and reasoning AI inference."

Jensen Huang, NVIDIA, GTC 2025 keynote

"The combination of NVIDIA Dynamo and NVIDIA GB200 NVL72 creates a powerful compounding effect that optimizes inference performance for AI factories that are deploying MoE models, such as DeepSeek R1 and the newly released Llama 4 models in production."

NVIDIA, on the inference rack + serving stack

"Oracle's state-of-the-art GB200 deployment... delivers exceptional performance and energy efficiency for agentic AI powered by advanced AI reasoning models."

NVIDIA & Oracle, on the first wave of Blackwell inference racks

The inference rack is real, shipping, and the unit of the market.

InferenceRacks.com is the literal phrase. Be the one who owns the address.

Acquire the domain →