Ethan Mollick:基准测试捕捉不到模型个性
中文全文 · AI 翻译
模型的个性很重要,而这种重要性是基准测试无法捕捉的。Opus 模型从大约 4.7 到 5 经历了一段不太好的时期,感觉不再那么“Claude”了,更像一个稀释版的 Fable(嗯,是教师模型的缘故吗?)。Opus 5.5 又让我找回了和老 Claude 一起工作的感觉。
对照原文
Model personality matters in a way that benchmarks can't capture. Opus models went through a rough patch from around 4.7 to 5 where they just didn't feel "Claude-y" anymore, more like an watered-down Fable (hmmm, teacher models?). Opus 5.5 feels like working with ol' Claude again