live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

vLLM 0.29 makes Model Runner V2 the default and removes deprecated paths

The release changes the default execution path, removes ten deprecated model architectures and ships new admission-control and model-support features.

Old vLLM path removed; Model Runner V2 becomes default.
AI-generated illustration
By The Release Desk· Sep 9, 2026the quick take — two AI hosts go live when you do

vLLM 0.29.0 changes the serving engine’s default execution path and removes several deprecated interfaces, making this more than a routine update for teams that package or operate vLLM-based inference services. The upstream release notes list 594 commits from 277 contributors.

What changed

Model Runner V2 is now the default for all models. The release adds CUDA-graph memory profiling for KV-cache sizing, batch-sharded sampling, prompt embeddings and additional speculative-decoding work to that execution path. Model Runner V1 remains in use for a limited set of ROCm models and features that V2 does not yet support.

The release also changes several defaults. FlashInfer all-reduce is enabled by default for tensor-parallel CUDA groups, while prefix-cache NONE_HASH behavior is deterministic by default. Operators gain --max-num-queued-reqs and --max-num-queued-tokens controls for admission management.

Model coverage expands to GraniteSWA and GraniteMoeSWA, alongside Hy4-preview, Qwen3.8-Flash-Next, NemotronH Omni Reasoning V3 and Kimi K3 NVFP4 checkpoints. The release includes additional work for Kimi K3, DeepSeek V4, speculative decoding, reinforcement-learning weight synchronization and Mamba prefix caching.

Who it affects

Platform teams should treat the update as a compatibility review rather than a drop-in patch. Ten deprecated model architectures have been removed. The PyAV video decoder backend is gone, and the python -m vllm.entrypoints.openai.api_server invocation is deprecated in favor of vllm serve. FlexOlmo, Olmo3 and Hunyuan V1/VL move to the Transformers modeling backend.

The default container and Python-package artifacts use CUDA 13.0. The project also publishes CUDA 12.9, ROCm, CPU and XPU artifacts, so image and accelerator choices need to remain explicit in deployment pipelines.

What to do

Before rollout, inventory workloads that rely on a removed architecture, PyAV decoding or the legacy Python module invocation. Test representative models against Model Runner V2, including memory sizing, speculative decoding and tensor-parallel behavior. CUDA operators that need the previous all-reduce path can opt out with VLLM_ALLREDUCE_USE_FLASHINFER=0.

Teams consuming vLLM through a downstream product should wait for that product’s validated build and support statement rather than substituting the upstream image directly. Teams operating upstream vLLM should pin the intended accelerator artifact and run compatibility and performance tests before promotion.

Filed by The Release Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.