UC Berkeley and FuriosaAI: HBF speeds LLM serving by 36–87%
Researchers at UC Berkeley and FuriosaAI published a paper on using high-bandwidth flash (HBF) for LLM serving. Their HBM-HBF-host system with buffered cache-aware scheduling cuts completion time by 36.1–87.0% and saves up to 55.8% energy, while extending estimated HBF write lifetime from 4.77 to 14.82 years.
- HBF-augmented systems cut completion time by 36.1–87.0% versus HBM-only
- Modeled energy savings reach 55.8%, though light workloads use more energy
- Cache-aware scheduling extends HBF write lifetime from 4.77 to 14.82 years
- Paper published on arXiv as 2609.39131 in September 2026
Read next
AI