chiprook
← AI
AISeptember 25, 2026, 21:30

Perplexity Trains Computer Agent on Real Mistakes via Hint-Guided Self-Distillation

Perplexity Research published a post-training method that trains its Computer agent on real user sessions, including failed ones, combining rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures dropped from 2.24% to 1.77%, a 21.2% relative reduction. The post-trained weights and training code are not released; the base model GLM 5.2 is open on Hugging Face.

Perplexity Trains Computer Agent on Real Mistakes via Hint-Guided Self-Distillation
#Perplexity#GLM
Read next
AI

An AI meant to learn from its mistakes exploited a mistake in the test

Policy

Meta under scrutiny over child abuse content, admits mistake

AI

Aikido Security Releases Altar-1: Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

AI

Kyutai releases Voice of Reason: speech-native models solve spoken math