llm-d 0.10.0 turns production hardening into an image and migration decision
The release separates signed production images from unsigned pull-request builds while deprecating several llm-d-maintained artifacts.
llm-d 0.10.0 is less about a single serving feature than about narrowing the project’s supported production path. The release notes frame the update around operational hardening and production readiness, then attach concrete migration work to those goals.
What changed
The project now separates its image pipeline by trust level. Official images built from merged code or releases remain on GHCR and are signed with Cosign, while images produced for open pull requests move to Quay and are intentionally unsigned. The release notes say this change currently applies only to inference-server images, not the router or sidecar images.
The release also moves more components toward upstream or consolidated homes. The llm-d-kv-cache repository is moving into llm-d-router, and the workload-variant-autoscaler repository has been renamed llm-d-autoscaling. Its WVA guides are deprecated and scheduled for removal in a future release.
Several artifacts now have explicit replacements. llm-d’s CUDA image is deprecated in favor of docker.io/vllm/vllm-openai, and its AWS image is deprecated in favor of AWS’s public vLLM image. The standalone filesystem connector is deprecated because its function is now available in vLLM’s in-tree OffloadingConnector. The latency-predictor component is also deprecated after the project concluded that its performance benefit did not justify its operational complexity.
Who it affects
Platform teams that pin llm-d image names or automate image admission need to distinguish GHCR production artifacts from unsigned Quay pull-request builds. Deployments using the project’s CUDA or AWS images, the old KV-cache repository, WVA guide paths, the standalone filesystem connector or the latency predictor all have migration work to assess.
The update also changes the component baseline. The release matrix moves the vLLM base image from 0.26.0 to 0.30.0 and updates router, inference-simulator, CPU, ROCm, XPU, async and batch-gateway components. llm-d cautions that some release-matrix tests could not be run because of infrastructure constraints, so the matrix should be checked per deployment path rather than treated as a blanket certification.
What to do
Before upgrading, operators should inventory image references and deprecated components against the 0.10.0 release matrix. Production policies should accept the signed GHCR artifacts and avoid treating unsigned Quay pull-request images as equivalent releases. Teams using the deprecated CUDA, AWS or filesystem-connector artifacts should validate the named upstream replacements, while WVA users should update repository and guide references to the autoscaling project.
Because the project reports incomplete testing for some paths, the practical upgrade gate is the matrix row for the hardware, model server and guide a deployment actually uses.
sources
- llm-d v0.10.0 releasegithub.com
comments · 0