live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

KServe 0.20 gives model-serving teams safer rollouts and a hardware-backed path for encrypted models

The release adds TEE-based model decryption, progressive traffic controls, CPU-backed KV-cache offload and first-class tracing for Kubernetes inference workloads.

KServe 0.20 adds declarative rollout, encryption, and tracing controls.
Side by side: what changed
By The News Desk· Sep 29, 2026the quick take — two AI hosts go live when you do

KServe 0.20 pulls several day-two model-serving controls into its Kubernetes APIs: encrypted-model handling inside trusted execution environments, progressive traffic management, CPU-backed KV-cache offload and declarative tracing. The project’s release post was published August 6.

What changed

For sensitive model artifacts, KServe can now download JWE-encrypted files and ask a local Confidential Data Hub inside a trusted guest to obtain the decryption key through hardware attestation. The model is decrypted inside the trusted execution environment rather than exposed to the host. The implementation works with both InferenceService and LLMInferenceService, and the project says it is compatible with Key Broker Service backends including Trustee and Intel Trust Authority.

The release also adds weighted traffic groups to LLMInferenceService. Operators can shift requests between service versions while readiness and degradation are tracked for each group independently. For conventional InferenceService workloads in RawDeployment mode, KServe 0.20 adds named canaries with percentage-based traffic allocation; promotion can retarget the stable service without recreating model pods.

For vLLM deployments, a structured kvCacheOffloading.cpu field now lets operators reserve CPU memory as a secondary KV-cache tier. KServe renders the corresponding vLLM transfer configuration, including for disaggregated prefill/decode layouts. A separate tracing API enables OpenTelemetry for the inference server and scheduler without requiring teams to inject environment variables through pod-template overrides.

Who should care

Platform teams already standardizing inference behind Kubernetes custom resources get the clearest benefit. The release moves rollout safety, model confidentiality and observability into declarative configuration instead of leaving each model team to assemble sidecars, environment variables and bespoke deployment logic.

There are useful runtime changes too. KServe now supports vLLM as a standalone InferenceService runtime, routes Anthropic Messages API traffic through the inference pool, and can mount OCI model images directly with Kubernetes ImageVolume. That native OCI path requires Kubernetes 1.33 or newer.

What to do

Teams testing 0.20 should start with one operational problem rather than enabling everything at once. Canary or weighted traffic controls are the lowest-risk place to validate rollback behavior. KV-cache offload should be benchmarked against latency and host-memory pressure under realistic prompts. Confidential serving needs a working attestation and key-broker chain, so it belongs in a threat-model review rather than a routine configuration toggle.

The release also upgrades its llm-d dependency to 0.8.0 and adopts llm-d.ai routing CRDs. Existing users should therefore review CRD and controller changes before upgrading production clusters.

sources

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.