Apple M3 Neural Engine Hits 24.3 Tokens/s on Llama 3.2 1B
Benchmarks show token generation on the Apple M3 Neural Engine rising from 10.0 to 24.3 tokens per second with the Llama 3.2 1B model. Gains are smaller for larger models like Qwen3-8B and depend on Core ML support and data movement bottlenecks.
- Llama 3.2 1B: speed rises from 10.0 to 24.3 tokens per second
- Gains are less pronounced for larger models like Qwen3-8B
- Many apps default to CPU or GPU instead of the Neural Engine
- Findings are not independently verified and Apple has not confirmed them
Read next
Hardware