Jerry Liu 介绍扫描表单的分层提取能力
我们已经非常擅长逐层提取扫描表单 📋——包括印刷字段、叠加文字、手写内容和审阅者的批注。
例如,对于带批注的许可证,我们的 Extract 代理既能提取原始许可证编号,也能提取其上绘制或手写的批注,并将它们准确对应到原文。
与使用前沿模型相比,我们的代理将错误率至少降低一半,而价格仅为其一小部分。
我们的 Extract v2.5 代理在性价比上达到领先水平,扫描表单只是其中一个例子。来试试吧!
Extract v2.5:https://www.llamaindex.ai/blog/introducing-extract-v2-5
可在 LlamaParse 中使用:https://cloud.llamaindex.ai/
对照原文
We got extremely good at extracting a scanned from layer by layer 📋 - from printed fields, overlaid text, handwriting, and reviewer annotations. For instance, on an annotated permit, our Extract agents can extract out both the original permit number but also any drawn/handwritten annotations on top. they're also grounded exactly in the source text. our agents reduce error rates by 2x+, at a small fraction of the price compared to using a frontier model. Our Extract v2.5 agents are SOTA in price/performance, and scanned forms are just one example. Come check it out! Extract v2.5: https://t.co/TTlTdECU03 Available in LlamaParse: https://t.co/XYZmx5TFz8