Notice Nearby

Company newsroom — sourced from NVIDIA · Notice Nearby

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

NVIDIA

September 16, 2026

Read original on the NVIDIA newsroom

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. Underlying all three is platform fungibility: the same infrastructure runs any model, any workload, from training to inference, recommender to reasoning, language to video, keeping utilization high. The NVIDIA platform is purpose-built to optimize across all these, as highlighted by MLPerf Inference v6.1 results released today: NVIDIA Vera Rubin NVL72 system debuts with leading performance: In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72. NVIDIA GB300 NVL72 scales with leading efficiency: A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline. Continuous software optimizations drive performance gains: Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0. Optimizations continued post-v6.1 submission, delivering further performance gains. For organizations making AI infrastructure decisions, performance, scaling efficiency and software velocity are important considerations that determine long-term inference economics. Vera Rubin NVL72 Makes MLPerf Inference Debut With Leading Performance NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL. Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72. These early results showcase NVIDIA’s accelerated pace of innovation and how performance will improve with continuous software optimizations. MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0106 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information. This performance means each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack, while lowering cost per token. The results reflect full-stack codesign across hardware and software. Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference, while NVFP4 precision reduces memory footprint across model weights, attention and KV cache — increasing throughput with minimal loss of output quality. Vera Rubin submissions heavily used disaggregated serving, separating prefill and decode along with large-scale expert parallelism for maximum efficiency across the mixture-of-experts layers that power models like DeepSeek-R1 and Qwen3-VL. The NVL72 scale-up domain — powered by sixth-generation NVIDIA NVLink and NVLink Switch to deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet — provides the interconnect foundation that makes these techniques effective at rack scale. This codesign extends to NVIDIA’s partner ecosystem: Nebius also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance. AI agents, which reason, plan and act across multiple steps, are reshaping how inference performance is measured. In benchmarks designed to capture this shift, such as SemiAnalysis AgentX, Vera Rubin NVL72 delivered 30 x better performance than GB300 NVL72 in preview testing. In addition, the upcoming MLPerf Endpoints benchmark will bring standardized measurement to agentic inference workloads, beyond what traditional throughput benchmarks capture. NVIDIA GB300 NVL72 Scales With Leading Efficiency Scaling efficiency — how effectively additional GPUs translate to throughput gains — is a key measure of AI infrastructure productivity. NVIDIA delivers this with high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks and efficient request orchestration across nodes. NVIDIA’s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario. Throughput grew nearly in proportion to the hardware added. MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information. Scaling efficiency is key because more GPUs don’t automatically mean proportionally more throughput. If adding nearly double the GPU count delivered only a single-digit percentage improvement in throughput, the infrastructure cost would far outpace the performance return. The architecture, interconnect and software must all scale together. GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video — 9x higher throughput and 7.5x lower latency than a single node. Software Optimizations Drive Continuous Gains NVIDIA platform undergoes continuous software development, delivering performance and feature improvements. In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results. The gains came through lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and NVIDIA Dynamo. Software optimization continued past the v6.1 submission deadline as well. Post-submission results, not yet verified by MLCommons, on GPT-OSS-120B and DLRMv3 show further performance gains. AI Inference at Every Scale Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B. The NVIDIA partner ecosystem participated broadly, with 19 partners — eight of them on multi-node Blackwell NVL72 systems — demonstrating excellent performance. This includes ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro and Wiwynn. From compact edge devices to the largest AI factories, NVIDIA continues to advance performance across the full technology stack with an annual cadence of platform architectures, continuously improving software and an ecosystem built to deliver it at scale. Learn more about the NVIDIA Vera Rubin platform. Categories: AI Infrastructure Hardware Networking Software Tags: Dynamo Inference NVIDIA Vera Rubin TensorRT Related News AI Infrastructure NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut Sep 16, 2026 AI Infrastructure Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers Sep 16, 2026 AI Infrastructure AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories Sep 15, 2026 AI Infrastructure Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now Aug 27, 2026 — Company newsroom — sourced from NVIDIA. Matter furnished by the company. Not a Notice Nearby paid placement. This page reprints matter furnished by the company from its official newsroom. Notice Nearby did not write this release. Read the original: https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference

Company newsroom — sourced from NVIDIA. Matter furnished by the company. Not a Notice Nearby paid placement. NN-PR-NR-2026-0294. This page reprints matter furnished by NVIDIA from its official newsroom. Notice Nearby did not write it, and it is not a $79 paid placement. It is not a legal public notice, not an obituary, and not an official Notice Nearby announcement. Record of Sale, LLC · Oregon.

The chain stores a hash, not the notice. The hash is not statutory publication. View on blockchain. Base stores the content hash, publication number, press-release id, and timestamp — not the full release text. The complete release remains on Notice Nearby. This is advertising, not a legal notice.

1d564938 691aac20 4657c220 8d9d30ed 9648d2de 4456a67d 9d1e68e2 89c99ae9

All press releases · Official newsroom

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut · Company newsroom — sourced from NVIDIA · Notice Nearby