Jerry Liu:用目标和评估定义智能体任务
如今,你几乎可以通过定义一项评估,并围绕它反复优化,来解决任何任务,而不必直接定义用于解决任务的确定性工作流或智能体工作流。
数据提供公司的全部工作,就是为所有经济活动定义评估,让前沿模型能够胜任任何事情。
这样一来,你的工作就变成了为前沿智能指明正确方向——定义要解决什么,以及如何衡量什么样的结果才算好。
智能体应用的界面将随之演进,以体现这一点。最复杂的流程仍然需要某种显式的工作流构建界面,但大多数任务都可以归结为目标和评估指令。
对照原文
These days, you can pretty much solve any task by defining an eval and hillclimbing over it, instead of directly defining the deterministic/agentic workflow to solve it. The data provider companies' entire job is to define evals for all economic activity to make the frontier models capable of doing anything. Your job then becomes pointing frontier intelligence in the right direction - defining what to solve, and how to measure what good looks like. Agent application interfaces will evolve to capture this. The most complex processes will still need some sort of explicit workflow builder interface, but most tasks can be compressed into goals and eval instructions.