vLLM’s LlavaOnevision2 loader bypassed its remote-code safety switch
The high-severity flaw affects vLLM before 0.28.0 and can execute code from a malicious model even when operators disable remote code.
vLLM has disclosed a high-severity arbitrary-code-execution flaw in its LlavaOnevision2 processor-loading path. The advisory affects versions earlier than 0.28.0 and lists 0.28.0 as the patched release. For teams that allow users or pipelines to select models, this is an upgrade decision rather than a theoretical implementation detail.
What failed
The affected loader calls Hugging Face Transformers’ get_class_from_dynamic_module() to import two processor classes from a model repository. vLLM passed trust_remote_code into that function, but the called function does not define that parameter. It absorbs the value through **kwargs and continues to cache and import the remote module.
The practical result is that setting trust_remote_code=False did not protect this path. A crafted LlavaOnevision2 model could place arbitrary Python in its processor module, and loading the model would execute that code with the authority of the vLLM process or container.
The advisory describes this as a guard asymmetry: vLLM’s safer wrapper resolves the remote-code decision before loading, while the LlavaOnevision2 implementation called the lower-level Transformers function directly. Because the architecture itself is built into vLLM, configuration loading could succeed without triggering an earlier remote-code gate.
Who should act
Operators running vLLM before 0.28.0 should prioritize the update if their serving environment can load models that are not fully controlled and reviewed. The advisory scores the issue 7.8 under CVSS 3.1, with a local attack vector and required user interaction, so it is not described as an unauthenticated network exploit. The risk appears when an operator, user or automation loads an attacker-controlled model while relying on trust_remote_code=False as a safety boundary.
Red Hat product status
Red Hat now rates the issue Important and marks 27 Red Hat AI Inference Server component builds as affected, spanning CPU, CUDA, Neuron, ROCm, Spyre, TPU and Gaudi variants. Its CVE record listed no issued errata when last modified on Sept. 12. Customers should therefore check Red Hat’s product record rather than assume the upstream 0.28.0 release is already present in their supported image stream.
What to do
Upgrade to vLLM 0.28.0 or later where upstream packages are used. Until an update is complete, restrict model sources to artifacts that have been reviewed and pinned, and do not treat trust_remote_code=False as sufficient protection for LlavaOnevision2 on affected versions. Teams should also inspect model-serving workflows for places where model identifiers can be supplied by tenants, users or external automation, because those are the paths where the faulty assumption matters most.
sources
- vLLM security advisory GHSA-3c86-2m5g-59q7github.com
- Red Hat CVE-2026-90553access.redhat.com
comments · 0