PhantomEnvironments: Training AI Agents in Fictional Worlds
Researchers propose PhantomEnvironments, a method for training LLM agents with reinforcement learning in synthetic fictional worlds generated by pure rule systems, no LLM required. Agents trained on entirely fictional data transfer search and reasoning skills to real benchmarks like HotpotQA, avoiding annotation costs and benchmark contamination.
- Worlds are generated by templates and rules without LLMs, at near-zero marginal cost
- Agents learn search strategy rather than facts and transfer it to real tasks
- Hop count drives transfer more than other environment features
- Limits include overfitting to templates and a narrow action space
Read next
AI