Novita open-sources Chord kernels for INT4 mixture-of-experts serving in vLLM
Chord adds shape-aware W4A16 CUDA paths for Kimi K2.x, while its grouped vLLM integration remains unfinished.
Novita AI has open-sourced Chord, a set of CUDA operators for serving INT4 mixture-of-experts models with BF16 activations, and connected its indexed execution path to vLLM’s existing Humming backend. In a September 15 engineering report published by vLLM, the developers describe kernels tuned around the routed-token shapes seen when serving Kimi K2.x models.
Two paths for different serving shapes
Chord ships two kernel families. Its indexed path consumes vLLM’s sorted expert-routing buffers and covers H200 expert-parallel prefill, H200 tensor-parallel serving, H200 decode and B200/B300 decode. Separate grouped contiguous and grouped masked paths target prefill and decode on Hopper-class GPUs with different routing and weight layouts, according to the project.
The design responds to a practical problem in mixture-of-experts serving: prefill and decode send very different numbers of tokens to each expert. Chord chooses schedules from the routed shape it receives rather than treating total token count as the only useful signal. The implementation uses different tile sizes, occupancy limits and pipeline depths across those regimes.
What the measurements show
Against the matching public Humming paths, the project reports per-layer gains of 1.11 to 1.20 times for H200 expert-parallel prefill and 1.16 to 1.24 times for H200 expert-parallel decode. In an earlier end-to-end Kimi K2.6 serving run on eight H200 GPUs, the indexed path improved combined prefill throughput by 9.6% and decode output throughput by 4.1% to 8%, depending on batch size, the report says.
The largest published figure, 2.15 times for B300 decode, needs a qualification: Chord was compared with Humming’s default untuned configuration because the public Humming project did not ship a tuning table for that GPU. The authors also say the measurements are kernel-level results, not a guarantee of equivalent end-to-end gains for every deployment.
What operators can use now
The indexed path can be installed as a Python package and selected in compatible vLLM revisions with the existing humming quantization backend. It supports the compressed-tensors INT4 group-32 checkpoint format used by Kimi K2.x and rejects unsupported schemes instead of silently choosing an incompatible kernel.
The grouped operators are less mature. Their standalone API is available, but integration with vLLM’s Humming backend is still work in progress. Platform teams evaluating Chord should therefore distinguish the usable indexed route from the grouped roadmap, then benchmark complete serving behavior—including routing, communication and activation costs—on their own model and concurrency mix.
sources
comments · 0