chiprook
← AI
AISeptember 29, 2026, 03:22

FlashAttention-4 ships for Blackwell B200 with FP4 and async pipelining

Tri Dao's lab released FlashAttention-4, an attention algorithm co-designed for NVIDIA Blackwell B200 and GB200. It adds asymmetric kernel pipelining via TMA, warp specialization and native FP4/MXFP4 execution, with reported throughput gains of 2.8–3.4x in FP8.

FlashAttention-4 ships for Blackwell B200 with FP4 and async pipelining
#Nvidia#FlashAttention#Blackwell
Read next
AI

"Drunk" AI models leak secrets and jailbreak more easily

AI

Kyutai releases Voice of Reason: speech-native models solve spoken math

Hardware

Nvidia B200 GPUs resell at 58% above launch price

AI

Nvidia details Task-Seeded SDG pipeline for Nemotron training