返回 X 名人动态

Aravind Srinivas 介绍让 Computer 从工具调用错误中学习的方法

中文全文 · AI 翻译

我们正在分享关于我们后训练方法的新研究,该方法教导 Perplexity Computer 智能体通过模仿良好轨迹并明确纠正可避免的错误(如错误的工具调用,即使整体轨迹是成功的)来从真实用户会话中学习。该方法结合了拒绝采样微调 (RFT) 与提示引导的自蒸馏,它在实时 A/B 测试中将工具调用失败率降低了约 21%。

引用 @perplexity_ai 的推文
对照原文

We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agent to learn from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls (even when the overall trajectory was successful). The method combines rejection sampling fine‑tuning (RFT) with hint‑guided self‑distillation, and it cuts tool call failures by about 21% in live A/B tests

老杨AI实操

微信扫一扫,添加好友

老杨AI实操的微信好友二维码

手机可长按保存图片,再到微信中识别二维码

保存二维码