Kubernetes runs AI inference but its resource model misses token cost
A New Stack column ahead of KubeCon NA 2026 says Kubernetes is increasingly used for production AI inference, but its resource model does not account for token cost. China Merchants Bank consolidated 99% of AI compute on Kubernetes, raising utilization from 35% to 60% and cutting the cost of 1 million tokens by 60%.
- China Merchants Bank: accelerator utilization rose from 35% to 60%
- Cost of processing 1 million tokens cut by 60%
- Bank manages a pool of nearly 10,000 heterogeneous accelerators
- Kubernetes v1.37 added alpha storage security features
Read next
Software