live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
newsAI

vLLM uses Mooncake to separate Kimi K3 serving from DSpark training

A new five-billion-parameter draft model and hidden-state connector turn speculative decoding into a capacity-planning choice across separate inference and training nodes.

Three nodes split serving and training, connected by hidden-state transfer.
AI-generated diagram
By The News Desk· Sep 17, 2026the quick take — two AI hosts go live when you do

The vLLM Project has extended its open-source Speculators training library to Kimi K3, a 2.8-trillion-parameter model, and released a five-billion-parameter DSpark draft model under the RedHatAI namespace. In the project’s September 15 engineering post, the team reports that the speculator raises single-stream generation on math workloads from about 110 to 435 tokens per second per user and delivers as much as 3.5 times more output throughput at matched interactivity under concurrent load.

What changed

DSpark uses a parallel draft-model backbone, then adds a Markov logit-bias head to restore local token dependencies and a confidence head to decide how much of a proposed block should be verified. The released Kimi K3 speculator proposes eight tokens per decoding step and reached a macro-average acceptance length of 4.11 tokens across nine evaluation domains, according to the vLLM results.

The more consequential engineering change is in training. Kimi K3 is too large for the project’s earlier single-node arrangement, even with four-bit weights. The team built a MooncakeHiddenStatesConnector that separates target-model serving from draft-model training and streams hidden states between vLLM and Speculators processes over RDMA or TCP.

The source supports a three-node working set: two GB300 nodes, each with four GPUs, serve the quantized Kimi K3 model while one four-GPU node trains the speculator. The team says that arrangement produced the best throughput among the configurations it tested. Mooncake lets operators scale the serving and training sides independently instead of forcing both into one fixed allocation.

Why platform teams should care

The work turns a research technique into a deployable artifact rather than only publishing benchmark numbers. vLLM provides a container recipe that loads RedHatAI/Kimi-K3-speculator.dspark directly through its speculative-decoding configuration. The same Speculators path has also been validated with draft models for Qwen3.6, Gemma 4 and GLM 5.2, the project says.

For teams operating large-model inference, the practical decision is whether the extra draft model and its training capacity are justified by workload shape. The gains are strongest in the project’s math and long-context tests, while acceptance length varies by domain. Operators should reproduce those measurements with their own prompts, concurrency and latency targets before planning capacity around the headline throughput increase.

The reusable pattern is the separation itself: vLLM serves the target model, Speculators trains the drafter, and Mooncake moves the required hidden states between them. That gives platform teams separate scaling controls for two workloads with different resource profiles. The published checkpoint offers a quick deployment path; the connector is the part that makes workload-specific retraining practical beyond one node.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.