NVIDIA's PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes
NVIDIA researchers with Princeton and the University of Maryland introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. It trains agents to avoid pivotal mistakes and recover when they happen, recovering from 72.7% of replayed pivotal mistakes versus 20.3% for standard OPD.
- PivotOPD ranks first on all 8 per-benchmark averages against 13 baselines
- Recovery from pivotal mistakes: 72.7% vs 20.3% for standard OPD
- 59% of failed Qwen3 rollouts contained a pivotal mistake, median turn 8–12 of 30
- On SWE-Bench Verified, a Nemotron-3.5-SFT student improved from 62.8% to 66.0%
Read next
AI