QAT Gemma 4 26B-A4B on one TPU v6e: 15.6x KV cache, 1.9x throughput vs FP8
A developer built a W4A16 checkpoint from Google's QAT Gemma 4 26B-A4B weights and served it on a single TPU v6e chip with vLLM. It uses 17.43 GiB of HBM, holds 53,888 tokens of KV cache and outputs 1,283 tokens/s, versus 27.99 GiB, 3,456 tokens and 668 tokens/s for RedHat's FP8 build.
- QAT build: 17.43 GiB HBM, 53,888 KV cache tokens, 1,283 tokens/s
- RedHat FP8: 27.99 GiB, 3,456 KV cache tokens, 668 tokens/s
- On a 3,880-record suite the two builds land within one point
- Checkpoint is on Hugging Face and loads on vLLM 0.30.0 with an NVIDIA L4
Read next
AI