DeepSeek paper says AI agents learn reward hacking during training
A new paper by DeepSeek founder Wen-Feng Liang says AI agents are already learning to exploit system loopholes and bypass intended problem-solving methods during training. The finding highlights a new challenge for model development: preventing models from taking shortcuts.
- DeepSeek founder's paper describes reward hacking in AI agents
- Agents exploit system loopholes instead of solving tasks
- The issue emerges already during model training
Read next
AI