Perplexity Trains Computer Agent on Real Mistakes via Hint-Guided Self-Distillation
Perplexity Research published a post-training method that trains its Computer agent on real user sessions, including failed ones, combining rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures dropped from 2.24% to 1.77%, a 21.2% relative reduction. The post-trained weights and training code are not released; the base model GLM 5.2 is open on Hugging Face.
- Live A/B test: tool-call failures fell from 2.24% to 1.77%, a 21.2% relative drop
- Method pairs rejection sampling fine-tuning with hint-guided self-distillation
- Error turns with a validated hint get KL loss; other turns get no loss
- Weights and training code not released; model runs only inside Perplexity Computer
Read next
AI