Test-Time Training Collapses When Agents Learn From Their Own Rollouts
A KAUST study (arXiv:2610.05076) found that test-time training collapses when an AI agent updates fast weights on its own generated tokens. On 760M TTT-E2E, self-generated updates added 6.00 nats of loss, and ALFWorld success dropped from 86.8% to 27.6%. The authors propose a Settlement buffer protocol.
- Self-training added 3.01 nats of loss at 125M and 6.00 nats at 760M
- ALFWorld success fell from 86.8% to 27.6% at update scale 8
- WebShop exact task success dropped from 16.0% to 10.5% across five seeds
- Settlement cut the loss gap to 0.07 nats at 125M and -0.02 at 760M
Read next
AI