GKE Pod Snapshots Cut Model Load Times by Up to 89%
Google published benchmarks for GKE Pod snapshots: a 70B model loads in 37 seconds and an 8B model in 15 seconds, cutting startup latency by as much as 89%. The feature saves a pod's running state, including CPU and GPU memory, and requires GKE Sandbox on gVisor.
- 70B model loads in 37 seconds, 8B in 15 seconds
- Startup latency reduced by up to 89%
- Requires GKE Sandbox on gVisor, cluster 1.35.3+
- Codeway cut Retake startup from a minute to 8 seconds
Read next
AI