AREX-2: Self-Improving LLM Agents via Long-Horizon Reflection
Researchers introduced AREX-2, a Qwen3.8-27B-based agent trained on verifiable self-improvement trajectories in ML engineering and programming. It scored 81.8 on MLE-bench Lite, 92.2 on GAIA and 93.8 on DeepSearchQA, with performance still improving after up to 30 refinement rounds.
- Built on a Qwen3.8-27B backbone trained on synthesized verifiable improvement trajectories
- Scores: 81.8 on MLE-bench Lite, 92.2 on GAIA, 93.8 on DeepSearchQA
- Accuracy keeps rising up to 30 reflection rounds, while GPT-4o plateaus after 5–10
- Reflective-Self-Correction loop distills state updates to limit context contamination
Read next
AI