vLLM brings day-zero serving support to NVIDIA Vera Rubin
The upstream project has daily CUDA 13.4 images, Rubin-tuned kernels and preliminary results showing as much as 7.84× the per-GPU AgentX throughput of GB200.
vLLM says its serving stack now runs on NVIDIA’s Vera Rubin NVL72 platform, with daily container builds, initial Rubin-tuned kernels and support for models including DeepSeek, Kimi, GLM and MiniMax. The work is an early hardware-enablement milestone rather than a finished performance release, but it gives inference teams a concrete path to start testing the next NVIDIA rack architecture.
What changed
The project’s first Rubin build is available as the vllm/vllm-openai:cu134-nightly image, using CUDA 13.4 and PyTorch 2.15. vLLM says Rubin remains compatible with many kernels built for NVIDIA’s Blackwell architecture family, while FlashInfer 0.7.0 adds Rubin-specific attention, GEMM and mixture-of-experts kernels.
A more structural optimization uses CUDA 13.4 locality domains to place shards of mixture-of-experts weights near the streaming multiprocessors that consume them. In vLLM’s preliminary MiniMax M3 layer tests, that placement produced about a 1.2× speedup for small-token forward passes. The project says the feature remains under active design and development.
The collaboration also gives Red Hat a direct role in the upstream enablement. vLLM credits NVIDIA and Red Hat with setting up the daily Rubin Docker builds, alongside contributions from Inferact, NVIDIA and the broader project community.
What the numbers mean
The headline result comes from SemiAnalysis AgentX: vLLM serving MiniMax M3 on Vera Rubin NVL72 reached up to 7.84× the throughput per GPU of NVIDIA GB200 at matched interactivity. Under a 150-tokens-per-second constraint, the reported advantage was 5.18×.
A separate MLPerf Inference 6.1 result used vLLM behind NVIDIA Dynamo to serve Qwen3-VL-235B-A22B. vLLM reports up to 3.7× the throughput of GB300 NVL72 across offline, server and interactive vision-language scenarios.
Those figures should be read as early, workload-specific results, not a blanket expectation for every model. vLLM explicitly calls the work an early look and lists substantial follow-on tasks, including fuller locality-domain support, additional FlashInfer integration and more Rubin-specific attention and MoE kernels.
Who should act
Teams planning Rubin evaluations can begin by reproducing their own model and latency targets with the CUDA 13.4 nightly image. The useful signal is not the largest multiplier by itself; it is that the upstream serving path, container pipeline and representative model coverage already exist before broader deployment.
Production teams should keep the nightly label in view. The sensible next step is controlled benchmarking against an existing GB200 or GB300 baseline, including interactivity constraints and the exact parallelism strategy, rather than treating the published peak as a capacity-planning constant.
comments · 0