China Telecom open-sources Xing4.0-29B-A4B agentic MoE trained on Ascend
China Telecom has released weights and configs for Xing4.0-29B-A4B on Hugging Face and ModelScope, an agentic Mixture-of-Experts model with 29B total and about 4B active parameters per token. It is the first model at this scale trained entirely on Ascend 910C NPUs with MindSpore, with vendor-reported training throughput gains of roughly 96%.
- 29B total parameters, ~4B active per token, 64 routed experts
- 256K native context extensible to 512K, 40 layers, MLA attention
- Trained on Ascend 910C clusters via MindSpore/MindFormers
- Serving via vLLM, SGLang, KTransformers; fine-tuning in LLaMA-Factory
Read next
AI