Huawei Debuts OceanStor M900: PB-Scale Shared KV Cache With ~60μs NPU-to-SSD Hop
Huawei unveiled OceanStor M900 Context Memory Storage at HUAWEI CONNECT 2026 in Shanghai, a storage layer built for AI inference SuperPoDs. A single cluster scales to 64 PB of KV cache, with NPU-to-SSD access latency of about 60 microseconds (a claimed 90% cut) and 40 TB/s aggregate bandwidth.
- Cluster capacity reaches 64 PB, lifting per-NPU KV cache from GB to TB
- NPU-to-SSD latency about 60μs, a claimed 90% reduction
- Aggregate access bandwidth of 40 TB/s, 1.5x peer solutions
- Up to 24 DWPD and a claimed 16x SSD endurance extension
Read next
Hardware