Xiaomi Claims 80% Compute Cost Cut for Unreleased MiMo V3
Xiaomi's HySparse2 design for the unreleased MiMo V3 model claims a 4.5x memory reduction and 80% lower compute costs. DeepSeek's V4.1 Flash uses layer note-sharing and 4x per-token memory compression; independent verification of long-context recall remains pending.
- HySparse2: 4.5x less memory, 80% lower compute costs
- DeepSeek V4.1 Flash cuts per-token memory 4x
- Xiaomi claims recall accuracy up to 256,000 tokens
- MiMo V3 is unreleased with no independent testing
Read next
AI