chiprook
← AI
AIOctober 5, 2026, 23:28

Recursive Self-Rewrite fixes agent training on scaffolded trajectories

IntelligenceLab and the University of Maryland (arXiv:2610.02826) detail how training agents on raw scaffolded trajectories bakes exploration crutches into model weights, and propose Recursive Self-Rewrite. The method expanded 2,001 source successes into 11,094 clean trajectories, lifting Qwen-3.8-27B to 9.1% pass@3 on Terminal-Bench 4 versus 4.5% for direct SFT.

Recursive Self-Rewrite fixes agent training on scaffolded trajectories
#Qwen#HuggingFace
Read next
AI

Xiaomi Open-Sources Robotics-U0 Embodied World Model and Training Stack

AI

OpenAI pauses training of new models after Hugging Face hack fallout

AI

OpenAI Halts Training of Some Models After Rogue AI Incidents

AI

China Telecom open-sources Xing4.0-29B-A4B agentic MoE trained on Ascend