NVIDIA Vera Rubin NVL72 Delivers Top Performance in MLPerf Inference v6.1

NVIDIA Vera Rubin and Blackwell Platforms Advance AI Inference Performance and Efficiency

NVIDIA System performance, efficient infrastructure scaling, and continuous software optimization are becoming increasingly important factors in determining the economics of AI inference. As organizations deploy AI applications at scale, the ability to generate more tokens, serve more users, and extract greater value from existing infrastructure can directly influence the overall cost and productivity of AI operations.

Higher system performance enables more tokens to be generated in a given amount of time, potentially allowing businesses to serve more users and increase the value generated from each AI system. Efficient scaling ensures that adding computing resources translates into proportional increases in throughput, helping organizations expand AI services without requiring excessive additional infrastructure. Continuous software optimization further improves the return on existing investments by allowing the same hardware to deliver increasingly higher performance over time.

Underlying these capabilities is platform flexibility. AI infrastructure must be capable of supporting different models and workloads, from training and inference to recommendation systems and advanced reasoning applications, as well as workloads spanning language, vision, and video. Infrastructure that can support a broad range of applications can maintain higher utilization as workloads evolve.

NVIDIA says its full-stack platform is designed to address these requirements, with the latest results from MLPerf Inference v6.1 highlighting advances in performance, scaling efficiency, and software optimization.

Vera Rubin NVL72 Makes Its MLPerf Debut

NVIDIA’s next-generation Vera Rubin NVL72 platform made its first appearance in MLPerf Inference with preview results on two demanding workloads: DeepSeek-R1 and Qwen3-VL.

According to NVIDIA, Vera Rubin NVL72 delivered up to 3.7 times higher throughput than GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios. The results used vLLM together with NVIDIA’s open-source Dynamo inference framework.

On DeepSeek-R1, Vera Rubin NVL72 achieved up to 2.5 times higher throughput than GB300 NVL72, with the results using NVIDIA TensorRT-LLM.

NVIDIA says these early results demonstrate the performance potential of the Vera Rubin architecture while also illustrating how ongoing software development can further improve results.

Higher throughput can allow each Vera Rubin NVL72 rack to generate more tokens and serve more users than a comparable GB300 NVL72 rack. For organizations operating AI services at scale, increased throughput can also contribute to lower cost per token and more efficient use of data center resources.

The results reflect a combination of hardware and software optimization. Vera Rubin’s enhanced Tensor Cores and Transformer Engine are designed to accelerate both the prefill and decode stages of inference, while NVFP4 precision reduces the memory requirements associated with model weights, attention, and KV cache.

The platform also makes extensive use of disaggregated serving, separating prefill and decode operations, as well as large-scale expert parallelism. These techniques are particularly relevant for mixture-of-experts models such as DeepSeek-R1 and Qwen3-VL.

Rack-Scale Networking Supports AI Performance

High-performance networking plays a central role in enabling these techniques. The Vera Rubin NVL72 scale-up domain uses sixth-generation NVIDIA NVLink and NVLink Switch technology.

NVIDIA says the interconnect provides significantly higher packet rates and lower latency compared with conventional Ethernet, helping GPUs communicate efficiently within the rack.

This tightly integrated combination of computing, memory, networking, and software is designed to enable AI workloads to operate efficiently across large numbers of GPUs.

NVIDIA’s partner ecosystem is also participating in the rollout. Nebius submitted Vera Rubin NVL72 preview results as part of MLPerf Inference v6.1, demonstrating performance on the new platform.

As AI applications evolve toward more autonomous systems, conventional throughput benchmarks are also being supplemented by workloads designed to measure agentic performance. NVIDIA says Vera Rubin NVL72 delivered 30 times higher performance than GB300 NVL72 in preview testing on SemiAnalysis AgentX, a benchmark focused on AI agents that reason, plan, and execute multiple steps.

The upcoming MLPerf Endpoints benchmark is expected to provide another standardized approach for evaluating agentic inference workloads, expanding performance measurement beyond traditional throughput-focused tests.

NVIDIA

GB300 NVL72 Demonstrates Efficient Scaling

While Vera Rubin highlights the performance potential of NVIDIA’s next-generation architecture, GB300 NVL72 demonstrated the importance of scaling efficiency in MLPerf Inference v6.1.

Scaling efficiency measures how effectively additional GPUs translate into increased throughput. Simply adding more accelerators does not guarantee proportional performance improvements. Communication overhead, networking limitations, software inefficiencies, and workload distribution can all reduce the benefits of additional hardware.

NVIDIA’s DeepSeek-R1 submission demonstrated 99% scaling efficiency when scaling from one GB300 NVL72 rack containing 72 GPUs to four racks containing 288 GPUs in the offline scenario.

According to NVIDIA, throughput increased nearly in proportion to the additional hardware.

Such efficiency is important for large-scale AI infrastructure because poor scaling can cause infrastructure costs to increase much faster than useful output. Maintaining high scaling efficiency requires the architecture, interconnect, networking, and software stack to work together.

GB300 NVL72 also demonstrated rack-scale performance on the WAN 2.2 text-to-video benchmark. NVIDIA reported throughput of 0.65 720p videos per second, with each video taking approximately 5.7 seconds. The company says this represented nine times higher throughput and 7.5 times lower latency compared with a single node.

Continuous Software Optimization

NVIDIA also highlighted the role of ongoing software development in improving AI infrastructure performance.

In MLPerf Inference v6.1, GB300 NVL72 performance on Qwen3-VL improved by up to 1.6 times compared with v6.0 results. NVIDIA attributed these gains to several software improvements, including lower KV

Source Link:https://blogs.nvidia.com/

Share your love