AWS Launches SageMaker HyperPod Inference Gateway for GPU-Aware Routing
AWS introduced Amazon SageMaker HyperPod Inference Gateway, a Kubernetes-native router for LLM inference that is aware of GPU state. According to AWS, time to first token drops by up to 82%, and P95/P99 latencies fall by 94–98% under mixed workloads.
- Installs with one command aws eks create-addon, no sidecars or code changes
- Time to first token reduced by up to 82%: from 4.4s to under 800ms
- P95/P99 dropped 94–98% in tests for Llama-3.1-70B and Qwen3-32B
- Tier 2 Global Inference Router with 35s cross-cluster failover coming soon
Read next
AI