live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
analysisAI

vLLM maps Qwen3.8 prefill/decode tradeoffs on GB300 NVL72

A reproducible benchmark reaches 5,000 total tokens per second per GPU at one end of the curve and 180 generated tokens per second per user at the other.

Benchmark frontier chart for vLLM Qwen3.8 tuning tradeoffs
Chart: figures from the story
By The News Desk· Sep 22, 2026the quick take — two AI hosts go live when you do

The vLLM project has published a reproducible tuning study for disaggregated serving of Qwen3.8-2.4T, giving platform teams something more useful than a single peak number: a map of the tradeoff between aggregate throughput and per-user interactivity.

In an 8,192-token input and 1,024-token output workload, the project reports reaching 5,000 total tokens per second per GPU in its throughput-oriented configuration and 180 generated tokens per second per user at the low-latency end of the Pareto frontier. Those are different operating points, not one configuration delivering both results at once.

Memory is the first constraint

The study starts by accounting for the model’s memory use rather than jumping directly to cluster-scale tuning. Qwen3.8-2.4T combines 69 gated-delta-network layers with 23 full-attention layers. Full-attention state grows with token count, while the recurrent state for the other layers is allocated per request.

That distinction determines vLLM’s KV-cache block size and, ultimately, the number of concurrent requests each decode engine can hold. The authors estimate that one request in the tested workload consumes about 759 MiB of KV cache. They also show that model weights are only part of the budget: CUDA context, NCCL buffers, activation peaks and reserved CUDA-graph memory leave materially less than the nominal 279 GB available on each GB300 GPU.

For operators, the practical point is that topology selection is a memory-allocation decision as much as a compute decision. In the project’s measurements, a TP4/DP4 decode topology left more room for KV cache than TP8 and raised the estimated per-engine request capacity.

Tune prefill and decode separately

The project measured prefill with an 8,192/2 workload and decode with a 1/1,000 workload before combining the strongest candidates into disaggregated deployments. Lower-concurrency prefill favored TP4/DP2 with expert parallelism, while TP2/DP4 became stronger as concurrency increased. On decode, multi-token prediction improved performance while enough KV-cache capacity remained, but different topologies took the lead at higher concurrency.

This decomposition is the most transferable part of the work. Teams can measure the two phases independently, identify their bottleneck and then decide how many prefill endpoints to place in front of a fixed decode topology instead of exhaustively testing every full deployment.

Reproducible, but narrowly scoped

The published environment is specific: an NVIDIA GB300 NVL72 cluster, an NVFP4 Qwen3.8 checkpoint, a nightly vLLM development revision, a development build of Dynamo, srt-slurm 1.0.98 and AIPerf 0.12.0. The project links the exact srt-slurm recipes and reports 95% on GSM8K for each selected prefill/decode configuration.

That makes the result inspectable, but it is not evidence that the same frontier holds on other accelerators, prompt distributions or service-level objectives. OpenShift AI teams using vLLM should treat the recipes as a tuning method and baseline—not as a production capacity promise.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.