DiffusionGemma turns one vLLM pass into a typed decision engine
A merged vLLM change lets DiffusionGemma fill fixed answer slots and expose confidence, while Red Hat maps a preview path onto OpenShift AI.
A merged vLLM change gives open-model teams a new serving primitive: ask DiffusionGemma several bounded questions, fill their answer slots in one denoising step and return probabilities that application code can act on. Red Hat has now published a deployment walkthrough and product roadmap for testing that pattern on Red Hat AI and OpenShift AI.
What changed in vLLM
The merged vLLM pull request adds the machinery for a “structured read.” A client seeds DiffusionGemma’s fixed-length token canvas with a response template, leaves only the answer positions unset and limits the request to one denoising step. vLLM then returns log probabilities for each allowed token; the client can select the top answer and derive uncertainty from the distribution.
The approach supports yes-or-no, scored and multiple-choice questions, and several questions can share one request. Choices must currently map to single tokens so that the fixed canvas does not shift. A client can translate longer labels such as moderation_spam to a token such as B and convert it back after inference.
This is not yet a standard vLLM endpoint. The pull request includes a prototype /v1/systemone server and low-level controls for seeded canvases, read-only requests and step limits. Red Hat’s walkthrough likewise labels the example server and custom runtime route as unsupported preview work.
The early performance signal
The pull-request author tested one DiffusionGemma deployment on a single NVIDIA DGX Spark. With a 32-token canvas and three decisions per request, it measured 8.7 requests per second at one-way concurrency and 54 requests per second at 32-way concurrency—about 162 decisions per second. Those are author-reported prototype results, not an independent benchmark, and the PR’s own task tests include misses on some classifications.
The operational trade-off is memory. Red Hat notes that diffusion state buffers grow with batch size, canvas length and the model vocabulary. Small decision canvases can sustain much higher concurrency than ordinary generation, but platform teams still need to size against their chosen model variant and schema.
Where Red Hat AI fits
Red Hat says DiffusionGemma 26B-A4B is already a validated model for its existing multimodal uses and that optimized FP8-dynamic and NVFP4 checkpoints are available in the RedHatAI collection. The company says the NVFP4 variant uses roughly one-third of the memory of BF16.
For now, teams can prototype the structured-decision mode with an upstream vLLM nightly and an unsupported custom serving runtime on Red Hat OpenShift AI, or use a Red Hat AI Inference preview image. Red Hat’s stated sequence is a preview after structured-read support reaches a stable vLLM release, followed by hardening based on feedback. Its post names an example-server Developer Preview for OpenShift AI 3.6 GA if vLLM 0.31 or later lands, while leaving the hardened endpoint’s timing and support level undecided.
The practical appeal is not another chatbot. It is a self-hosted component for routing, moderation, risk scoring and agent branching where applications need fixed answers and explicit confidence. The merged code establishes the primitive; supportability, interface stability and workload-specific accuracy are the next gates.
sources
- Run decision models on vLLM and Red Hat AI using DiffusionGemmadevelopers.redhat.com
- vLLM PR #57250: structured generation mode for DiffusionGemmagithub.com
comments · 0