live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
analysisAI

Red Hat’s MLPerf 6.1 results make topology—not Kubernetes—the tuning problem

OpenShift led the four-GPU GB200 field, while CPU-only vLLM results exposed memory bandwidth and scheduler stability as the practical limits.

OpenShift leads GPU benchmarks while CPU-only results reveal topology limits.
Chart: figures from the story
By The News Desk· Sep 29, 2026the quick take — two AI hosts go live when you do

Red Hat’s MLPerf Inference 6.1 submissions put OpenShift at the top of the four-GPU NVIDIA GB200 field and show the same vLLM engine producing valid results on CPU-only Intel systems. The more useful finding is what made the difference: NUMA placement, memory bandwidth and scheduler isolation rather than a new model-serving abstraction.

MLCommons says Inference 6.1 drew a record 30 participating organizations and added end-to-end RAG and edge-agentic tests. Red Hat’s submission concentrated on existing datacenter workloads across gpt-oss-120b, Qwen3-VL, Llama 3.1 and Whisper.

OpenShift reaches the front of the GB200 group

A four-GPU Supermicro GB200 NVL4 system running OpenShift 4.21 and vLLM 0.24 produced 61,205.6 tokens per second in the offline gpt-oss-120b test and 51,095.2 tokens per second in the server scenario. Red Hat reports both as the highest results among four-GPU GB200 submissions, with its offline result 20% above the only other system at that scale.

Normalized per GPU, Red Hat calculates roughly 15,300 tokens per second—higher than the larger GB200 NVL72 submissions in this round. That supports a narrower conclusion than “Kubernetes has no overhead”: a tuned OpenShift system can lead this specific peer-reviewed configuration set.

The tuning detail is more transferable. On the Qwen3-VL workload, unpinned vLLM workers placed most memory on the remote socket and reduced host-to-device bandwidth from about 166 GB/s to about 50 GB/s. Pinning workers per socket restored bandwidth and improved end-to-end throughput by roughly 9%.

CPUs remain viable for bounded workloads

Red Hat and Intel also ran vLLM on a two-socket Xeon 6972P system with no accelerators. It delivered 1,507.21 tokens per second offline for Llama 3.1 8B and 1,966.11 samples per second for Whisper Large v3. Red Hat reports the highest per-core throughput among two-socket CPU-only submissions for both workloads.

The team reserved one CPU core per rank for serving and benchmark overhead, limited deep idle states and found memory bandwidth—not core count—to be the first-order limit. Those choices point to CPU inference for moderate-throughput, latency-tolerant work, not as a general substitute for accelerators.

What teams should test

Platform engineers should reproduce topology and scheduler settings before comparing hardware prices. A useful evaluation should record NUMA placement, memory bandwidth, reserved cores, idle-state policy and tail-latency behavior—not just aggregate token throughput.

The broader result is software portability: vLLM appeared in 22 submitters’ stacks, and Red Hat used it across both GB200 GPUs and Xeon CPUs. The benchmark does not prove every workload will move unchanged between them, but it does show that inference-engine investment can span substantially different hardware classes.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.