Red Hat validates 64 MI300X GPU slices on OpenStack Services on OpenShift
Eight MI300X GPUs exposed 64 schedulable SR-IOV virtual functions, while Red Hat documented performance, topology and production-support boundaries.
Red Hat has validated a high-density AMD Instinct MI300X virtualization design for Red Hat OpenStack Services on OpenShift, showing that eight physical accelerators can become 64 schedulable SR-IOV virtual functions for tenant virtual machines.
What changed
The test system ran Red Hat OpenStack Services on OpenShift 18.0 FR5 with RHEL 9.6 compute nodes. AMD GPU-IOV Manager placed each MI300X in CPX mode, exposing eight virtual functions per device. Nova and Placement inventoried those VFs as PCI resources, while VFIO and libvirt attached them directly to guests.
Red Hat exercised the maximum-density layout of 64 VMs with one VF each. It also passed configurations with eight, 16 and as many as 32 VFs assigned to one VM; a single VM with all 64 VFs remains under investigation because of virtual-IOMMU behavior at very high PCI passthrough counts.
For the primary BabelStream operations, the virtualized configurations delivered about 96% to 99% of the bare-metal memory-bandwidth baseline. Repeated lifecycle runs also passed: the 64-VM layout completed 100 boot-and-shutdown cycles, with VF allocation and reclamation verified and no VF leaks observed in completed tests.
Who it affects
The design targets OpenStack platform teams that want better utilization and hardware-backed isolation for independent AI and HPC workloads. Each CPX VF exposes about 24 GB of VRAM, making the pattern a fit for inference services, embedding generation, retrieval-augmented generation and isolated research jobs that do not need an entire 192 GB MI300X.
The boundary is just as important as the density result. In the tested configuration, IOMMU isolation blocked direct peer-to-peer DMA between VFs. CPU-to-GPU transfers remained available, but tightly coupled multi-GPU applications that depend on high-bandwidth P2P communication may need physical-GPU assignment or bare metal. Placement also matters: VFs sharing one physical GPU contend for its resources.
What to do
Red Hat points operators to the MI300X Validated Architecture in the OpenStack K8s Operators repository and recommends beginning with one GPU-backed VM before scaling. Configure Nova with the VF product ID 74b5, not the physical-function ID, and validate the application's memory, topology and communication requirements.
Do not turn the lab procedure into a production shortcut. The validation used a lab-built DKMS GIM package, and Red Hat says installing GIM with DKMS directly on EDPM hosts is not a supported production workflow. Production deployments should use the documented Red Hat and AMD driver-delivery and support path.
sources
comments · 0