Aravind Srinivas 介绍让 Computer 从工具调用错误中学习的方法
中文全文 · AI 翻译
我们正在分享关于我们后训练方法的新研究,该方法教导 Perplexity Computer 智能体通过模仿良好轨迹并明确纠正可避免的错误(如错误的工具调用,即使整体轨迹是成功的)来从真实用户会话中学习。该方法结合了拒绝采样微调 (RFT) 与提示引导的自蒸馏,它在实时 A/B 测试中将工具调用失败率降低了约 21%。
对照原文
We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agent to learn from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls (even when the overall trajectory was successful). The method combines rejection sampling fine‑tuning (RFT) with hint‑guided self‑distillation, and it cuts tool call failures by about 21% in live A/B tests