AI 财务建议评测:方法和评估者同样重要
我认为,我们确实有很多充分理由担心 AI 的答案是否正确,尤其是免费模型。但这与近期一些认为 AI 已能很好地提供财务建议的研究相矛盾,所以我很好奇这里采用了什么研究方法。我下载了报告;它来自一家销售不同 AI 金融服务产品的公司。
报告提供的信息有限,只列举了少数问题示例,因此无法判断其准确性。不过,我想知道,有没有熟悉英国税法的人能判断:GPT-6 Pro 对 Haiku 答案的辩护是对的,还是 Saturn 的批评是对的?
这一切都是为了说明,我认为政府需要做得更好,直接对模型进行评估并公布结果,而不是简单地信任那些有自己议程的第三方评估者。了解人工智能能力实际发生了什么,似乎真的非常重要。
对照原文
I think there are lots of good reasons to worry about whether AI is giving correct answers (especially free models), but, given the contradiction with other recent research suggesting AI is now doing well at financial advice, I was curious about what the methodology was here. I downloaded the report, which is from a company selling different AI financial services products. Given the limited information in the report, it is impossible to judge it's accuracy, since there are only limited examples of questions. However, I am curious if anyone familiar with UK tax law knows whether GPT-6 Pro's defense of Haiku's answer is right, or if Saturn's critique is.
All this is to say that I think governments need to do a better job doing direct evaluations of models and publishing results rather than simply trusting third-party evaluators with their agendas. It seems really important to know what’s actually happening with AI ability.

