Edge0 Framework Runs Qwen 35B AI Model from Your SSD
The open-source Edge0 framework from The Stack allows running a 35-billion-parameter MoE model by streaming 92.9% of weights from storage and using only 2.9 GB of active memory. It requires an SSD with up to 4 GB/s throughput, and model accuracy drops by 3.9 points due to reducing active experts from eight to four.
- Active memory reduced to 2.9 GB by streaming 92.9% of weights from storage
- Requires storage with up to 4 GB/s throughput
- Active experts per layer reduced from 8 to 4, accuracy drops 3.9 points
- Plans include 1-bit versions and support for complex tasks
Read next
AI