FlashAttention-4 ships for Blackwell B200 with FP4 and async pipelining
Tri Dao's lab released FlashAttention-4, an attention algorithm co-designed for NVIDIA Blackwell B200 and GB200. It adds asymmetric kernel pipelining via TMA, warp specialization and native FP4/MXFP4 execution, with reported throughput gains of 2.8–3.4x in FP8.
- Asymmetric pipeline: producer warps load via TMA, consumer warps compute on Tensor Cores
- Native FP4 and MXFP4 support on B200's 20 PFLOPS and 8 TB/s HBM3e
- Reported 2.8–3.4x throughput gain in FP8
- Benchmarks at 32k, 128k and 1M token contexts vs FlashAttention-3 and PagedAttention
Read next
AI