ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction· ExtractBench:企业文档提取的基准
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We r
企业文档提取的首个全面基准测试平台ExtractBench
- 核心方法
- 基于4,869页、370份文档、8个业务领域和67种文档类型的测试集,结合独立系统一致性、已知值和人工验证
- 适合谁读
- 适合研究者、工程师和产品经理阅读
- 要解决的问题
- 评估企业文档提取系统的性能,特别是方案指导下的准确性和完整性
- 关键实验
- 包含4,869页文档,370份企业文件,8个业务领域,67种文档类型,明确了不同挑战场景的标签
- 主要贡献
- 提供了一个全面评估企业文档提取系统的新基准,涵盖准确性、完整性和成本
- 意义与局限
- 帮助企业优化文档处理流程,提高自动化水平;但局限于特定领域和文档类型