live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

vLLM 0.31 turns fast restarts and request trust into upgrade decisions

The release adds GPU-resident weight preloading and broader distributed-serving controls while changing multimodal request handling and several operator-facing flags.

Before and after vLLM 0.31 upgrade changes
Timeline: dates from the story
By The Release Desk· Oct 5, 2026the quick take — two AI hosts go live when you do

vLLM 0.31.0 adds a GPU-resident restart path, expands large-scale serving controls, and changes several request and configuration defaults. The combination makes this more than a routine package refresh for platform teams operating shared inference services. The upstream release was published October 5, 2026.

What changed

The new vllm preload command starts a weight-cache daemon that keeps post-quantized model weights in GPU memory across engine restarts. The release extends that path to data parallelism and multi-token-prediction draft models, and adds health and readiness checks. An experimental snapshot workflow also uses CRIU to restore a fully initialized tensor-parallel-one engine.

Distributed-serving changes include a MoonEP all-to-all backend, prefill context parallelism with data parallelism, DeepEPv2 sequence parallelism, and back-pressure detection for KV-cache offloading. The scheduler also gains --max-num-active-seqs, separating active admission from the existing maximum sequence setting.

The security boundary for multimodal requests is stricter. Per-request mm_processor_kwargs and media_io_kwargs are rejected unless the server starts with --trust-request-mm-kwargs. The release also changes prefix-cache hashing to distinguish extra-key sources and include the LoRA path.

Who it affects

Operators that restart engines during model rollouts can evaluate preloading as a way to avoid reloading post-quantized weights. Teams running disaggregated or expert-parallel serving have new controls, but should benchmark them against their current topology rather than assume an automatic gain.

API owners need to review any client that sends per-request multimodal processor or media I/O options. Those requests now fail by default unless the deployment deliberately opts into trusting them.

What to do

Test 0.31.0 in staging with production request shapes and the same accelerator topology used in service. Audit clients for the newly gated multimodal fields before rollout.

The upgrade also requires a configuration review: tokenizer_mode="slow" is removed; --enable-mamba-fine-grained-prefix-cache is renamed to --enable-mamba-shared-prefix-checkpoint; XPU graphs are enabled by default; and several quantization and backend options changed or disappeared. Pin the release, compare startup and steady-state behavior, and update manifests before promoting it into a shared serving environment.

Filed by The Release Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.