live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
guideAI

Kueue and DRA make OpenShift GPU quotas reflect the hardware workloads consume

Red Hat build of Kueue 1.4 maps device classes to queue quotas and can account for partitioned GPUs by memory rather than raw device count.

Full GPU versus partitioned slices under different quota accounting.
Side by side: what changed
By The News Desk· Oct 2, 2026the quick take — two AI hosts go live when you do

Dynamic Resource Allocation gives Kubernetes a richer description of accelerators, but it does not decide how scarce devices should be divided among teams. Red Hat build of Kueue 1.4 fills that policy gap on OpenShift by bringing DRA requests into Kueue's quota, borrowing, preemption and fair-sharing machinery.

That matters because the traditional device plug-in interface reduces every accelerator to an integer. A full 80 GB GPU and a small partition can each appear as one device, even though their capacity and scheduling value differ sharply. DRA exposes attributes such as memory, topology and compute capability; Kueue uses those declarations to make admission decisions before jobs reach the scheduler.

Whole-device accounting

Administrators map a DRA DeviceClass, such as gpu.nvidia.com, to a logical Kueue quota resource. A workload then references that class through a ResourceClaimTemplate, optionally adding a CEL selector for a specific product. Kueue resolves the class, charges the requested device count against the ClusterQueue, and suspends the workload when capacity is unavailable.

Once the DRA resource is represented in the queue, existing controls apply. Teams can borrow idle accelerator quota from another queue in the same cohort, higher-priority work can preempt lower-priority jobs, and admission fair sharing can distribute access among tenants. Red Hat describes whole-device quota support as production-ready.

Account for partitions by memory

Counter-based quota addresses partitioned devices. Instead of charging every full GPU or MIG slice as one unit, administrators configure a counter source such as GPU memory and express queue capacity in memory units. Kueue reads the counter exposed for the selected device or partition and charges the workload proportionally—for example, roughly 80 GB for a whole device versus 10 GB for a smaller profile.

On OpenShift, the operator detects the Kubernetes partitionable-device capability and enables the corresponding integration when counter sources are configured. Platform teams should validate the counter names and attributes exposed by their DRA driver before building quotas around them.

What to plan for next

Several improvements described in the article are upstream or future work, not capabilities to assume in Red Hat build of Kueue 1.4. Upstream Kueue 0.19 enables partitionable-device support by default. Extended-resource integration aims to preserve familiar resources.requests syntax while unifying quota with DRA claims. Consumable capacity targets dynamic sharing such as time-slicing and MPS, while scheduler-library integration could test whether admitted workloads are actually placeable on real nodes.

For a rollout, begin with a single device class and queue cohort, verify accounting for both successful and suspended jobs, and test borrowing and preemption deliberately. Add counter-based memory quota only after confirming that full devices and partitions report consistent units. Kueue can enforce fair admission, but administrators still need monitoring for jobs that hold quota while pods remain unschedulable.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.