Why Vera Rubin Wins: NVFP4 Precision Meets Purpose-Built Silicon
Vera Rubin NVL72's throughput lead over GB300 NVL72 traces to a combination of hardware and software changes rather than any single trick. NVIDIA describes the system's enhanced Tensor Cores and Transformer Engine as accelerating both the prefill and decode stages of inference, while NVFP4 precision shrinks the memory footprint of model weights, attention, and KV cache - raising throughput with what the company calls minimal loss of output quality [1]. Those software gains ride on a rack that packs 72 Rubin GPUs and 36 Vera CPUs with 20.7 TB of HBM4 GPU memory, 1,400 TB/s of GPU memory bandwidth, and 216 TB/s of NVLink bandwidth, rated at 3,600 PFLOPS of NVFP4 inference compute [2]. The DeepSeek-R1 result (2.5x) ran on TensorRT-LLM across offline, server, and interactive scenarios, while the larger Qwen3-VL gain (3.7x) used vLLM paired with NVIDIA Dynamo - a reminder that the advertised leap is really two separate stacks tuned for different workload shapes, not one universal multiplier.


