Study Tests 7 Attention Mechanisms on Latin Square in 49-Layer Model
VIDRAFT AI Research (arXiv:2609.20269) tested seven attention mechanisms arranged on a Latin square in the 49-layer Aether-7B-5Attn model (6.59B parameters MoE, ~2.98B active). Layer order was irrelevant; only having a mechanism from a different family helped: removing Mamba-2 worsened CE by 2.14%, while sliding window, differential, and full attention could be removed without effect.
- Aether-7B-5Attn: 6.59B parameters MoE, 49 layers, 144.2B tokens, ~11,700 B200-hours
- Removing Mamba-2 from four mechanisms gives +2.14% CE, other three within noise
- Homogeneous stack loses +1.68% at 700.9M and +2.63% at 1.514B parameters
- Significance threshold |Δ| > 2·pooled SD fixed before final runs
Read next
AI