Azure Unveils AI‑Optimized Kubernetes Service with OpenAI Integration
Microsoft released Azure AI-Optimized Kubernetes on September 10, 2026: OpenAI models run as sidecar containers in a managed cluster. It claims up to 45% lower inference latency, pricing from $0.00012 per token, and public preview in the US and Europe.
- OpenAI models run as sidecar containers without separate inference services
- Inference latency reduced by up to 45% per Microsoft internal benchmarks
- Pricing from $0.00012 per inference token, spot VMs offer up to 70% discount
- Public preview in US and Europe, global launch in Q1 2027
Read next
Software