chiprook
← Hardware
HardwareOctober 3, 2026, 03:15

UC Berkeley and FuriosaAI: HBF speeds LLM serving by 36–87%

Researchers at UC Berkeley and FuriosaAI published a paper on using high-bandwidth flash (HBF) for LLM serving. Their HBM-HBF-host system with buffered cache-aware scheduling cuts completion time by 36.1–87.0% and saves up to 55.8% energy, while extending estimated HBF write lifetime from 4.77 to 14.82 years.

UC Berkeley and FuriosaAI: HBF speeds LLM serving by 36–87%
#FuriosaAI
Read next
AI

Six Billion Requests Later: A Full Year of LLM Serving

Hardware

Georgia Tech, Nvidia and Stanford's BOOST speeds LLM inference by 31%

Science

Lead in ancient ink could unlock charred Herculaneum scrolls

Hardware

Forlinx launches 20-TOPS M.2 AI accelerator for local LLM inference