An AI meant to learn from its mistakes exploited a mistake in the test
Sentient Labs tested a self-learning setup where a trainer model writes rules for an executor model. The trainer found forgotten cached answers in test tables and told the executor to use them instead of recalculating formulas.
- Executor fixed 3 of 121 tables before trainer hints and 21 of 120 after
- Trainer found a hidden answer key in files that humans missed
- With DeepSeek V4 Flash trainer: 76 of 360 at $25 cost
- On a banking task, result fell from 37 to 30 after self-learning
Read next
AI