chiprook
← AI
AIOctober 10, 2026, 07:36

vllm-mlx on Apple Silicon: batching gives up to 3.39x, memory is the limit

vllm-mlx 0.4.1 is an MLX-based inference server for Apple Silicon exposing OpenAI- and Anthropic-compatible APIs with continuous batching, paged KV cache and prefix caching. On an M4 Max with 128 GB, five concurrent requests raised aggregate throughput by 1.64x to 3.39x depending on the model, while unified memory remains the hard ceiling.

vllm-mlx on Apple Silicon: batching gives up to 3.39x, memory is the limit
#Apple#Vllm-mlx#MLX
Read next
Hardware

Atomically thin transistors surpass silicon's switching limit

Hardware

Power, Memory and Packaging, Not Transistors, Now Limit AI Chips

Hardware

AI-defined vehicles hit compute, memory and validation limits

Hardware

Samsung bets on zHBM as 2.5D memory nears its limits