chiprook
← AI
AIOctober 2, 2026, 00:15

AREX-2: Self-Improving LLM Agents via Long-Horizon Reflection

Researchers introduced AREX-2, a Qwen3.8-27B-based agent trained on verifiable self-improvement trajectories in ML engineering and programming. It scored 81.8 on MLE-bench Lite, 92.2 on GAIA and 93.8 on DeepSearchQA, with performance still improving after up to 30 refinement rounds.

AREX-2: Self-Improving LLM Agents via Long-Horizon Reflection
#AREX-2#Qwen
Read next
AI

REAL-Q uses dynamic gradient descent to fix LLM quantization

AI

Nokia open-sources AnyJev: a training-free layer that turns any open LLM into a calibrated decision model

AI

PrismML hopes its tiny LLM could change how we all use AI

AI

Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight