chiprook
← Hardware
HardwareSeptember 24, 2026, 16:13

Apple M3 Neural Engine Hits 24.3 Tokens/s on Llama 3.2 1B

Benchmarks show token generation on the Apple M3 Neural Engine rising from 10.0 to 24.3 tokens per second with the Llama 3.2 1B model. Gains are smaller for larger models like Qwen3-8B and depend on Core ML support and data movement bottlenecks.

Apple M3 Neural Engine Hits 24.3 Tokens/s on Llama 3.2 1B
#Apple#M3#CoreML
Read next
Hardware

2026 flagship chips compared: A20 Pro leads single-core, XRING O3 tops multi-core

Hardware

CXMT gains ground in DRAM as Samsung, SK Hynix shift capacity to HBM

Hardware

XFEON unveils Physical-AI matrix: Xinghe S1 edge chip, AGLobe brain, Xinghui Token workstations

Hardware

Imec and TSMC expand cloud access to advanced-node chip design