Four preflight gates for Red Hat AI Inference benchmarks on AMD GPUs
A reproducible Podman run starts by proving host drivers, rootless device access, endpoint health and workload comparability—not by collecting a headline throughput number.
A useful AMD GPU benchmark begins before GuideLLM sends its first request. Red Hat’s reference workflow for Red Hat AI Inference 3.5 is a single-host, rootless Podman setup, so it can answer whether one model-serving configuration behaves reproducibly on one machine. It does not by itself validate an OpenShift architecture, compare vendors or predict a distributed production service (Red Hat Emerging Technologies).
Platform and performance teams can turn the procedure into four explicit go/no-go gates.
Gate 1: prove the host sees the accelerator
Before starting containers, verify that amd-smi lists the GPU, the amdgpu kernel module is loaded, rocminfo reports the expected architecture, and /dev/kfd plus /dev/dri/render* exist. A failure here is a host or ROCm problem, not a vLLM performance result. The guide installs Podman, crun, curl, jq and pciutils with elevated privileges, but runs the inference and load-generation containers without sudo (Red Hat Emerging Technologies).
Gate 2: prove rootless GPU access
The user must belong to the video and render groups, and Podman must use the crun OCI runtime because the launch path relies on --group-add=keep-groups. The Red Hat AI Inference container receives /dev/kfd and /dev/dri, disables SELinux labeling for those device mounts, and runs a small PyTorch check to count visible GPUs. Do not proceed if that container-level count differs from the hardware you intended to test (Red Hat Emerging Technologies).
This gate is also where teams should record the exact immutable inputs: Red Hat’s guide pins the model, tensor-parallel size and full image tags for both the ROCm vLLM server and GuideLLM. Changing any of them creates a different experiment (Red Hat Emerging Technologies).
Gate 3: prove the endpoint before measuring it
The reference launches Qwen2.5-7B-Instruct on 127.0.0.1:8000, waits for the health endpoint, then sends a chat-completions request to the OpenAI-compatible API. Keep that functional check separate from the benchmark. A failed response, model download problem or startup timeout is a deployment failure; folding it into throughput data hides the cause (Red Hat Emerging Technologies).
Gate 4: match the workload to the question
Red Hat supplies three distinct test shapes. The balanced synthetic run uses 1,000 input and 1,000 requested output tokens across concurrency levels from one to 650. A four-turn profile repeats prior conversation context to exercise automatic prefix caching. A Mooncake trace replay adds request timing, token lengths and shared-prefix behavior closer to real traffic. These workloads are not interchangeable: the first maps saturation, the second emphasizes cache reuse, and the third tests a trace’s particular traffic distribution (Red Hat Emerging Technologies).
Store JSON and CSV output with a timestamped run identifier and the model, image tags, GPU count and tensor-parallel setting. Compare runs only when those inputs and the workload definition are controlled. The guide’s 600-second duration and 5% maximum error-rate constraint are part of the method, not incidental defaults (Red Hat Emerging Technologies).
The result is a defensible local baseline. It can identify promising model, GPU and serving settings before a clustered test, but production sizing still requires the network, scheduler, storage, replica and failure behavior that this single-host loop intentionally leaves out.
sources
comments · 0