GAGAR: New Reward Method for Code-Writing RL Agents
Researchers introduced GAGAR, a groupwise agentic grading and advantage redistribution method for reinforcement learning of code agents. Instead of a binary test pass-fail reward, an agentic grader ranks passing patches by minimal blast radius and cleanliness, while advantages are rescaled to preserve their sum. It was tested on MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T) models.
- Binary test rewards can't tell an 8-line fix from a bloated 40-line patch
- GAGAR ranks test-passing trajectories by minimality and code cleanliness
- Sum-preserving advantage redistribution avoids skewing other task domains
- Tests on MiMo-V2.6-Flash 310B and MiMo-V2.6-Pro 1.02T improved accuracy
Read next
AI