vllm-mlx on Apple Silicon: batching gives up to 3.39x, memory is the limit
vllm-mlx 0.4.1 is an MLX-based inference server for Apple Silicon exposing OpenAI- and Anthropic-compatible APIs with continuous batching, paged KV cache and prefix caching. On an M4 Max with 128 GB, five concurrent requests raised aggregate throughput by 1.64x to 3.39x depending on the model, while unified memory remains the hard ceiling.
- Qwen3-0.6B-8bit: 328.1 to 1,111.8 tokens/s (3.39x) at 5 requests
- Qwen3-30B-A3B-4bit: 98.1 to 233.3 tokens/s (2.38x)
- Qwen2.5-1.5B-Instruct-4bit: 196.9 to 322.2 tokens/s (1.64x)
- Speedup depends on model, architecture and memory bandwidth, not a fixed multiplier
Read next
Hardware