返回 X 名人动态

Jerry Liu 分享 16 款视觉语言模型的文档解析评测

中文全文 · AI 翻译

我们全面评估了 16 个最近的前沿视觉语言模型(VLMs)——包括 Opus 5.5 和 GPT-6 Sol/Luna ——在更高推理强度是否会导致更好的文档解析性能方面。

更高推理强度通常会在其他基准测试(编码、知识工作)上带来改进,但直到最近,它并不明显地原生提升了阅读 PDF 的能力。

结果: ✅ 在前沿模型中,Opus 5.5 相对于其价格具有最佳性能。它尤其擅长解析表格。 ✅ Astra 也相当不错,但起价比 Opus 更贵。 ✅ GPT-6 Luna 在文档解析的低价端更具吸引力

如果您要大规模解析文档,您仍然会想要一个专用的 OCR 解决方案,比如 LlamaParse (https://cloud.llamaindex.ai/),它以更低的价格提供更好的性能。

但如果您在像 Codex/Claude Code这样的应用中“在智能体循环内”解析文档,并且太懒得集成专用解决方案,那么 Opus 5.5 目前是领先者。 ParseBench 上的完整结果:

https://www.parsebench.ai/

Jerry Liu 的原推文配图 1
原推文配图 · 查看来源
对照原文

We comprehensively evaluated 16 recent frontier VLMs - including Opus 5.5 and GPT-6 Sol/Luna - on whether higher effort led to better document parsing performance. Higher effort typically leads to improvements on other benchmarks (coding, knowledge work), but up until recently it wasn't obvious that this natively improved capabilities for reading PDFs. Results: ✅ Out of the frontier models, Opus 5.5 has the best performance relative to its price. It's especially good at parsing tables. ✅ Astra is also quite good, but starts at a more expensive price than Opus. ✅ GPT-6 Luna is more compelling at the cheaper end of doc parsing If you're parsing documents at scale, you'll still want a dedicated OCR solution like LlamaParse (https://t.co/XYZmx5TFz8) that has better performance at a cheaper price. But if you're parsing docs "in the agent loop" within an app like Codex/Claude code, and you're too lazy to integrate a dedicated solution, then Opus 5.5 is currently the leader. Full results on ParseBench: https://t.co/PWczfhp0OX

老杨AI实操

微信扫一扫,添加好友

老杨AI实操的微信好友二维码

手机可长按保存图片,再到微信中识别二维码

保存二维码