Thomas Wolf:作弊骤降可能源于模型识别出了评估
中文全文 · AI 翻译
人们之所以担心,是因为作弊行为突然大幅减少,最可能的解释是“评估意识”:最新的 Opus 模型或许已经足够聪明,能够识别出这个基准是在测试作弊行为,并相应地调整表现。
如果是这样,这个基准就不再衡量模型“自然”的作弊倾向了。
对照原文
People are worried because the most likely explanation for such a sudden drop in cheating is evaluation awareness: the latest Opus models may now be smart enough to recognize that this benchmark tests for cheating, and behave accordingly. If so, the benchmark no longer measures the models' "natural" tendency to cheat.