chiprook
← AI
AISeptember 27, 2026, 05:35

QAT Gemma 4 26B-A4B on one TPU v6e: 15.6x KV cache, 1.9x throughput vs FP8

A developer built a W4A16 checkpoint from Google's QAT Gemma 4 26B-A4B weights and served it on a single TPU v6e chip with vLLM. It uses 17.43 GiB of HBM, holds 53,888 tokens of KV cache and outputs 1,283 tokens/s, versus 27.99 GiB, 3,456 tokens and 668 tokens/s for RedHat's FP8 build.

QAT Gemma 4 26B-A4B on one TPU v6e: 15.6x KV cache, 1.9x throughput vs FP8
#Google#Gemma#VLLM#TPU
Read next
AI

Gemma 4 on SageMaker: QAT weights decode 2.05x faster than bf16 on one L4

AI

Google releases Gemma 4 open model family in five sizes

AI

Qwen3.8-27B on one RTX 3090 vs two: +20% decode, +14% cold prefill

AI

TaichuAI Open-Sources ZDTaichu5.0-9B Spatial Multimodal Model