Jerry Liu 发布 OpenDocRouter:用统一 API 接入文档解析模型
今天我很兴奋地介绍 OpenDocRouter —— 一个用于文档解析的统一 API,集成了最新的前沿模型和开放权重模型。
有很多视觉语言模型(VLM)和 OCR 模型可用于文档解析:我们在 ParseBench 上有130 多个模型,而 HuggingFace 上搜索“ocr”会得到数千个结果。
在 OCR 供应商之间选择非常耗时。你需要弄清楚合适的提示词、处理速率限制、管理部署、与你使用的所有模型进行集成,并随着新模型的发布进行基准测试。
OpenDocRouter 提供了一套全面、透明的模型,处于价格与性能的最优权衡范围。它管理一个统一的 API,将文档转录为 Markdown。它以“成本价”提供所有前沿和开放权重模型,仅收取少量交易费用。它处理所有模型的速率限制,确保你可以处理海量流量。它甚至提供边界框和布局作为服务,这样你就可以为使用的任何模型添加与原文位置的对应关系(grounding)。
当新的 OCR 候选模型发布时,我们将在 ParseBench 上对其进行基准测试,并立即将其添加到 OpenDocRouter。
我们正在非常快速地添加更多模型,同时也在添加一些极其令人兴奋的功能改进(例如自动路由),目前正在推进。
我们欢迎你的反馈!
你可能会问,为什么我们要这样做,因为我们在 LlamaParse(https://cloud.llamaindex.ai/)中的主要重点是构建文档解析和提取的最佳前沿模型。
虽然我们相信我们的模型在各自的价格点上精度属于一流,但我们也承认它们无法覆盖所有可能的价格点;在低端,有一些极其廉价的模型(mineru、luna、paddleocr),在高端,可能有一些边缘案例你想要运行最新的前沿模型。
对照原文
Today I’m excited to introduce OpenDocRouter - a unified API for document parsing with the latest frontier and open-weight models. There are a lot of VLMs and OCR models that can be used for document parsing: we have over 130+ models on ParseBench, and a HuggingFace search for “ocr” turns up thousands of results. It’s extremely time consuming to choose between OCR vendors. You need to figure out the right prompts, handle rate limits, manage deployments, integrations with all models you’re using, and benchmark new models as they come out. OpenDocRouter provides a comprehensive, transparent set of models along the price-performance frontier. It manages a unified API to transcribe documents to markdown. It serves all frontier and open-weight models “at-cost”, with a small transaction cut. It handles rate limits with all models to ensure you can put massive volume through. It even offers bounding boxes and layout as a service, so that you can add grounding to any model that you’re using. When new OCR candidate models ship, we will benchmark them on ParseBench and immediately add them to OpenDocRouter. We’re adding a lot more models very quickly, and also adding some extremely exciting feature improvements (e.g. automated routing) as we speak. We welcome your feedback! Check it out: https://t.co/v8PieWWUNV Blog: https://t.co/ilVqqJRacP
You may ask why we’re doing this given that our main focus within LlamaParse (https://t.co/XYZmx5TFz8) is building the best frontier models for document parsing and extraction. While we believe our models are best in class in accuracy at their price points, we acknowledge that they don’t cover every possible price point; at the low end there are extremely cheap models (mineru, luna, paddleocr), at the high end there may be an edge case or two you want to run the latest frontier model over.