NPU thermal throttling in Copilot+ laptops: 45 TOPS only in short bursts
A deep-dive shows the advertised 45 TOPS of NPUs in Copilot+ laptops holds only under short bursts: during continuous local inference the chip exceeds 85 °C and generation drops from 25 to 12 tokens per second. The heat also accelerates battery wear and raises LPDDR5X memory latency.
- Advertised 45 TOPS NPU performance lasts only the first minutes
- Sustained inference drops speed from 25 to 12 tokens per second
- Chip temperature exceeds 85 °C, triggering thermal throttling
- Liquid metal repaste lowers temperatures by only 2 °C
Read next
Hardware