vllm-metal turns Apple Silicon into a concurrent local serving target
The vLLM plugin adds paged KV caching, packed attention and an OpenAI-compatible server for multi-request inference on Macs.
vLLM has made Apple Silicon a first-class target for concurrent local inference through vllm-metal, a plugin that combines vLLM’s serving machinery with Apple’s MLX and Metal execution stack. The project’s September 22 engineering post describes v0.28.0 as its first official release; the current Homebrew installation path provides v0.29.0.
What changed
The plugin reuses vLLM’s V1 scheduler, paged key-value cache management, chunked prefill, sampling and OpenAI-compatible frontend. Model layers come from mlx_lm, while vllm-metal replaces attention with a paged variable-length Metal kernel that preserves request boundaries inside a packed batch.
That division matters for concurrent workloads. Rather than padding every prompt to the longest request in a batch, vllm-metal concatenates scheduled query tokens and uses sequence offsets to keep requests separate. Its fixed-size KV pages let admitted requests grow without repeatedly reshaping a contiguous cache. A memory guard also reserves headroom for macOS and other applications, then sizes a fixed KV cache after a startup warmup.
The first official release added batched multi-token prediction, GGUF checkpoints, hybrid-attention model support and an accelerated prefill path for M5 systems. The documented feature set also includes LoRA adapters, structured outputs, pipeline parallelism across Macs, and experimental vision-language, embedding, reranking and speech-to-text paths.
Who should care
This is primarily a developer and workstation-serving release, not a replacement for clustered production inference. It gives teams using Macs an OpenAI-compatible endpoint for local coding agents, experiments and multi-user test workloads while retaining vLLM’s admission control and batching model.
The project reports lower time to first token and end-to-end latency than the compared engines at concurrency levels two and four for a 4-bit Qwen3.8-27B workload on an M5 Pro with 64 GB of memory. Those are project-run benchmark results rather than an independent evaluation, and the post notes that results cover completed requests using engine-specific 4-bit conversions.
What to try
On macOS 15 or later, the documented path is to add the project’s Homebrew tap, install vllm-metal, and launch a model with vllm serve. Existing clients can then use the local OpenAI-compatible /v1 endpoint.
Teams should start with an explicit --gpu-memory-utilization budget and reproduce the project’s benchmark against their own prompt lengths and concurrency. Several paths remain bounded: batched MTP is currently limited to Gemma 4 with greedy sampling and synchronous scheduling, while prefix reuse for Qwen hybrid models is experimental and cannot yet be combined with speculative decoding.
comments · 0