chiprook
← AI
AISeptember 26, 2026, 04:55

Gemma 4 on SageMaker: QAT weights decode 2.05x faster than bf16 on one L4

A test on an Amazon SageMaker ml.g6.xlarge endpoint with one NVIDIA L4 showed the Gemma 4 E2B QAT checkpoint (4-bit weights) decodes at 105.1 tok/s versus 51.3 for bf16, and serves 1077.25 tok/s at 16 parallel requests versus 619.1. Both scored the same on 40 checked questions, while 4-bit weights save 18% of GPU memory.

Gemma 4 on SageMaker: QAT weights decode 2.05x faster than bf16 on one L4
#Google#Gemma#Amazon#Nvidia
Read next
AI

Google releases Gemma 4 open model family in five sizes

AI

Google Research Introduces Retrieve-for-Train (R4T): RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out

AI

The Biological Computing Co. partners with AWS to sell neuron-derived AI video model

AI

AWS Launches SageMaker HyperPod Inference Gateway for GPU-Aware Routing