Red Hat and PyTorch test Helion as a tuned vLLM linear backend
A Red Hat–PyTorch engineering project reports double-digit throughput gains in some Hopper workloads, but the implementation remains in a fork while maintainers weigh upstream tradeoffs.
Red Hat and PyTorch engineers have integrated the Helion kernel language with vLLM’s linear-backend interface, aiming to replace several hand-specialized matrix-multiplication paths with one implementation selected by per-shape autotuning. The engineering report, published October 2, says the current code is available in a Red Hat fork rather than the main vLLM repository.
One kernel, several execution choices
The prototype exposes standard GEMM, Split-K and Swap-AB as tunable choices inside a single Helion implementation. An ahead-of-time tuner selects both the algorithm and lower-level configuration for each workload shape, instead of relying only on manually written dispatch heuristics.
For the current evaluation, the team targeted FP8 and INT8 quantized linear layers on Nvidia Hopper GPUs. Its hybrid dispatcher uses Helion under CUDA Graph replay for token counts up to 32, then falls back to vLLM’s default CUTLASS or DeepGEMM path for larger shapes. The authors say that boundary avoids CPU launch overhead and limits how many shapes need expensive tuning.
The published results cover several dense Qwen models and three 8-bit quantization formats on one H100 80GB GPU. At kernel level, the reported geometric-mean speedups range from 1.110× to 1.178× against CUTLASS, FlashInfer or DeepGEMM, depending on format. End-to-end ShareGPT serving tests showed gains across the evaluated combinations and exceeded 10% throughput improvement in some workloads. Those figures are project results from a defined benchmark setup, not a general performance guarantee.
The upstream question is operational
The implementation is described as production-ready in the team’s vLLM fork, which also carries tuning tools and pre-generated configurations. It has not yet landed upstream. The engineering post identifies the main obstacle as maintenance: broad model coverage would require a large set of pre-tuned configuration files that are hard to validate continuously.
The proposed compromise is to keep the Helion integration and a functional default configuration upstream while asking operators to tune for their own models and hardware before deployment. That shifts work to serving teams, and the authors acknowledge additional costs: fine-grained tuning can take hours, cold starts can trigger JIT compilation, and execution outside CUDA Graphs can erase gains through CPU overhead.
This matters for Red Hat’s AI stack because vLLM is a central model-serving component, while the project was supported by Red Hat’s OCTO and vLLM teams. The immediate result is not a new supported OpenShift AI feature; it is an upstream-facing experiment in making inference kernels easier to port and specialize without giving up performance. Future work named by the team includes Nvidia Blackwell, AMD GPUs, TPUs and mixture-of-experts backends.
sources
comments · 0