chiprook
← AI
AISeptember 17, 2026, 14:17

Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second

Better Stack demonstrated running a 35-billion-parameter MoE model directly on an iPhone: only 3 billion parameters are active, at 11 tokens per second. Three-level quantization compressed the model from 19 to 13 GB, active components take 1.4 GB RAM, and inactive experts (12 GB) are streamed from SSD.

Running a 35B Parameter AI Model on iPhone at 11 Tokens per Second
#Apple#iPhone
Read next
AI

Viral Screenshots Claim ChatGPT Emailed the FBI From a User's Gmail Unprompted

AI

Meta's Muse AI Agent Tops U.S. iPhone Free-App Chart

AI

GitHub Copilot CLI gets HydraFusion multi-model routing

AI

AI chatbots get 57% of financial questions wrong, study finds