live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

vLLM’s TT plugin brings Tenstorrent accelerators behind the standard serving API

The out-of-tree backend preserves vLLM’s OpenAI-compatible surface while adapting scheduling, parallelism and sampling to Tenstorrent mesh hardware.

vLLM API on one side, Tenstorrent mesh backend on the other.
AI-generated diagram
By The News Desk· Sep 7, 2026the quick take — two AI hosts go live when you do

vLLM has introduced an out-of-tree platform plugin for Tenstorrent accelerators, adding another hardware backend without putting Tenstorrent-specific code into vLLM core. The plugin keeps vLLM’s existing OpenAI-compatible API and request format, while moving hardware-specific scheduling, model registration and execution into the separately installed TT package.

What changed

The plugin automatically registers Tenstorrent hardware when the ttnn runtime from TT-Metal is available. It currently maps a range of text and multimodal model families—including Llama, Qwen, Mistral, Gemma, DeepSeek V3 and GPT-OSS—to Tenstorrent implementations shipped with TT-Metal. Administrators can also register model bundles from an external directory without editing the plugin source.

That compatibility layer does not pretend Tenstorrent devices behave like GPUs. Tenstorrent systems compile and trace programs for a mesh of cores and chips, so the plugin replaces conventional tensor- and pipeline-parallel ranks with a mesh configuration selected through MESH_DEVICE. The published implementation rejects standard -tp and -pp settings rather than silently accepting options it cannot honor in the usual way.

The scheduler also separates prefill-only and decode-only steps. vLLM’s upstream scheduler can mix prefill and decode work within a token budget, but the Tenstorrent path favors homogeneous, shape-stable batches that can replay a captured trace. Long prompts can still use chunked prefill, with decode steps interleaved so active requests continue making progress.

Who it affects

The immediate audience is teams evaluating Tenstorrent hardware for vLLM-based serving. Existing clients can keep the familiar API surface, but operators need to account for different topology and scheduling behavior. On 32-chip Galaxy systems, some models use one compiled program spanning the mesh and expose multiple data-parallel KV-cache lanes inside a single engine process rather than assigning a process to each rank.

The plugin can sample tokens on-device when a request fits that path. Requests needing log probabilities, penalties, masks or custom logits processing fall back to vLLM’s host-side sampler automatically. An asynchronous decode path overlaps host readback with later scheduling, but the project describes it as a steady-state optimization rather than a universal asynchronous execution model.

What to do

Prospective users should start with the plugin’s supported-model table and a validated vLLM release, then choose a mesh configuration for the target model instead of carrying over GPU rank settings. Workloads that rely on mixed prefill/decode behavior, speculative decoding or custom sampling should be tested explicitly: the current plugin does not support speculative decoding, and some sampling features leave the device fast path.

The broader engineering signal is the extension boundary. Tenstorrent’s backend changes core execution assumptions—batch shape, parallel topology and where sampling runs—yet the integration remains outside vLLM core. That makes the plugin a useful test of whether vLLM’s hardware interfaces can accommodate architectures that are not merely GPU-compatible accelerators.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.