Local AI runs on a 2020 Apple Watch Series 6, beating Raspberry Pi
Better Stack ran the 90M-parameter Falcon H1 LLM on a 2020 Apple Watch Series 6 with 1GB RAM at 15-24 tokens per second. The team adapted llama.cpp to bypass Core ML incompatibility, while a larger 135M-parameter model proved too heavy for the watch.
- Falcon H1 with 90M parameters runs at 15-24 tokens per second on the 2020 watch
- Token generation is 50 times faster than a first-generation Raspberry Pi
- Core ML cannot handle the Mamba 2 architecture, so llama.cpp was adapted
- A 135M-parameter model exceeded the watch's 1GB RAM and 32-bit addressing
Read next
AI