Xiaomi details HySparse2 for MiMo-V3: two-level KV sharing cuts 1M-token prefill FLOPs about 5x
Xiaomi's LLM-Core researchers released HySparse2, a hybrid sparse attention architecture with two-level KV sharing for long-horizon MiMo-V3-class agent workloads, in an arXiv paper. On matched 80B-A3B MoE models it lifts MRCR-v2 and RULER-v2 by 11.30 and 19.81 points over HySparse, and at 1M tokens cuts prefill FLOPs by 2.92x and 5.02x versus HySparse and Hybrid SWA.
- HySparse2 splits the backbone into a YOCO-style self-decoder and cross-decoder
- KV Bridging links only full-attention layers; KV Reuse selects at token level
- At 1M tokens prefill FLOPs drop 2.92x and 5.02x vs HySparse and Hybrid SWA
- FP8 KV-cache shrinks to about 2.69 GB vs 6.72 GB and 12.09 GB for baselines
Read next
AI