Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips
Z.ai described production inference for GLM-5.3-Flash on a cluster of over 100,000 Chinese AI accelerators. An AI agent based on the model did much of the work, tripling throughput with preparation in under two weeks.
- Cluster of over 100,000 Chinese AI accelerators, a record scale for such chips
- GLM-5.3-Flash: 320B parameters, 18B active, 1M-token context
- Throughput tripled, production preparation took under two weeks
- KDA kernel fixes merged upstream into Flash Linear Attention on August 27, 2026
Read next
AI