Google Cloud's RRSI uses regularization to stop agent harnesses overfitting benchmarks
Google Cloud AI Research published RRSI, a method for regularized recursive self-improvement of LLM agent harnesses. It constrains prompt and logic edits, screens out benchmark-specific leakage and unjustified complexity: gains up to 14.1 points on the evolution split and up to 4.7 points on five out-of-distribution benchmarks while using 30% fewer policy tokens.
- RRSI adds regularization to the agent harness self-improvement loop
- Gains up to 14.1 points on evolution set and 4.7 on OOD benchmarks
- Regularized harnesses used 30% fewer policy tokens
- Tested on Claude Opus 4.8 and Gemini 3.5 Flash across eight benchmarks
Read next
AI