By Michal Sutter
Publication Date: 2026-09-25 14:30:00
Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% relative reduction.
Is it deployable? Not directly. Perplexity has not released the post-trained weights or training code. The model runs only as a model option inside Perplexity Computer. The base model, GLM 5.2, is openly available on Hugging Face.
Why Outcome-Only Filtering Falls Short
Standard rejection sampling fine-tuning (RFT) judges each session and imitates only the successful ones. A successful outcome does not mean every step was correct. An agent can recover from a bad tool call and still deliver the right answer. Imitating that full trajectory can reinforce the error. Discarding failed sessions also throws away clear evidence of avoidable mistakes.
Imitate, Correct, or Keep as Context
Perplexity team separates 2 decisions: which sessions hold behavior worth imitating, and which turns hold mistakes worth correcting.
Each assistant turn gets 1 of 3 treatments:
- Imitate: non-error turns in successful sessions receive cross-entropy (CE) loss.
- Correct: error turns with a validated hint receive Kullback-Leibler (KL) divergence loss, in any…

