CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases· 企业级 Q&A 基准测试
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands
评估大规模企业文档集的LLM Q&A性能的新基准。
- 核心方法
- 构建了一个名为CorporateBench的人类验证的多任务Q&A基准测试集,包含了超过23万文档的评价语料库,通过四个合成的企业模型(员工数从12到10,000)来评估LLM在信息抽取和知识库查询两方面的性能。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 现有的LLM评估方法过于简单或不能使用企业内部数据,无法真实反映企业在使用LLM时的情况。
- 关键实验
- 关键实验包括对企业规模不同、文档数量不同的四家合成企业进行Q&A任务评估。
- 主要贡献
- 提供了接近真实企业规模的、跨文档逻辑一致性保证的评估基准,涵盖多任务和动态知识库特性。
- 意义与局限
- 该基准有助于更准确地评估和提升LLM在企业应用场景中的性能,但其基于合成数据的局限性仍需注意。