返回 X 名人动态

Jerry Liu:前沿模型最具有挑战性的任务之一,就是能够从文档中极其密集的表格里提取数千个值。

中文全文 · AI 翻译

前沿模型最具有挑战性的任务之一,就是能够从文档中极其密集的表格里提取数千个值。

我们的新 Extract v2.5 代理能够在长列表提取上达到 93%-96%+ 的准确率,包括那些跨越页面的单元格。

相比之下,astra 在这个基准上的准确率只有约 30%。

在这些情况下,前沿 VLMs 可能会提前停止或遗漏重复记录。

它们还难以将每个值追溯回来源。而我们能够将每一个提取的值都追溯回来源。

看看下面的视频、博客文章,以及 llamaparse!

博客:https://www.llamaindex.ai/blog/introducing-extract-v2-5

llamaparse:

https://cloud.llamaindex.ai/

引用 @jerryjliu0 的推文
对照原文

one of the most challenging tasks for frontier models is being able to extract thousands of values from extremely dense tables in documents. our new Extract v2.5 agents are able to get 93%-96%+ on long-list extraction, including cells that fall in between pages. in contrast, astra gets ~30% accuracy over this benchmark. in these cases, frontier VLMs can stop early or miss repeated records. they also struggle to attribute each value back to the source. we are able to attribute every single extracted value back to the source. check out the video below, the blog post, and llamaparse! blog: https://t.co/TTlTdECU03 llamaparse: https://t.co/XYZmx5TFz8

老杨AI实操

微信扫一扫,添加好友

老杨AI实操的微信好友二维码

手机可长按保存图片,再到微信中识别二维码

保存二维码