live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
analysisAI

vLLM turns DeepSeek-V4.1-Flash’s day-zero path into a 5.3× throughput gain

SWA bounded replay, CUDA graphs and model-specific kernel fusion reshape the latency-throughput trade-off for long-horizon agent serving.

Chart of throughput and latency gains from vLLM optimizations.
Chart: figures from the story
By The News Desk· Oct 8, 2026the quick take — two AI hosts go live when you do

Three weeks after DeepSeek-V4.1-Flash arrived, the vLLM community says it has pushed the model to 1.9× its day-zero speed at low concurrency and 5.3× its throughput under a 150-tokens-per-second constraint. The result matters less as a single benchmark number than as a map of where agentic inference is spending time: sliding-window cache state, kernel launches and memory traffic rather than one dominant matrix multiplication.

What changed

The central change is sliding-window-attention bounded replay. DeepSeek-V4.1-Flash keeps a compact global KV cache alongside a larger, uncompressed sliding-window cache for the last 128 positions in each of 40 layers. Instead of storing the sliding-window state at every prefix-cache boundary, vLLM now rebuilds the last 128 tokens after a cache hit. On the decoder side, it runs the upper layers only on those final positions rather than across the whole prompt.

That replay is not bit-exact, but vLLM reports no meaningful accuracy difference on GSM8K and GPQA within about 1.5 standard errors. It is enabled by default for the model and can be disabled with --no-swa-bounded-replay.

The runtime also captures the trimmed upper layers in separate CUDA graphs. Without that step, launch overhead can erase the gain on short prompts; with it, vLLM measured a 30%–40% reduction in prefill computation time.

The rest is in the kernels

vLLM integrated DeepSeek’s MegaAttention, Mega-mHC, Mega-Gate and DeepSelect work, then fused more of the remaining path. The changes include scoring only fixed sparse-attention candidates, overlapping mHC coefficient work on a side CUDA stream, and keeping intermediate values on-chip in a fused low-latency kernel.

For the KV cache, MegaAttention reads an NVFP4 format that vLLM says is 45% smaller than its previous FP8 cache. At very long contexts, sparse MQA kernels are reported as 14×–23× faster per layer at 512K tokens than scoring the whole context and masking unused positions.

Who should care

Teams evaluating DeepSeek-V4.1-Flash for coding agents or other long-running tool loops should re-benchmark rather than carrying forward day-zero capacity estimates. vLLM’s low-latency setup used four-way tensor parallelism with FlashInfer attention; its throughput setup instead used data-parallel attention with experts split across GPUs to avoid duplicating the model’s shared KV latent.

Those are architecture choices, not drop-in promises for every cluster. The published figures come from the SemiAnalysis AgentX workload and selected NVIDIA Blackwell configurations. Operators should reproduce the latency-throughput curve with their own prompt lengths, concurrency and service-level target.

The practical takeaway is that DeepSeek-V4.1-Flash no longer has one obvious deployment shape in vLLM. Low-concurrency latency and high-concurrency throughput favor different attention and parallelism choices, while bounded replay changes both prefix-cache storage and prefill cost. Capacity planning should test both ends of that curve.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.