InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries· InsufficiencyBench:评估大模型处理不完整法律咨询查询
Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially determine the legal outcome. We introduce InsufficiencyBench, the first legal benchmark targeting query-side insufficiency: whether a model recognizes when a query lacks legally material information, identifies what is missing, and refrains from premature conclusions. We formalize a taxonomy of eight canonical missing-element categories across three structural failure modes---switch, gating, and fatal prerequisite--- and construct 202 benchmark items (58 base queries, 144 deficient variants) spanning six legal domains and 24 US jurisdictions and annotated by practising attorneys. Evaluating ten frontier models, we fin
评估大模型在处理不完整的法律咨询时的表现
- 核心方法
- 引入 InsufficiencyBench 基准,分类八种缺失元素,涵盖三种结构失败模式,构建 202 个基准条目
- 适合谁读
- 研究者、工程师、法律工作者
- 要解决的问题
- 现有法律 AI 系统评估基准假设用户查询是完全明确的,但实际中用户经常遗漏关键事实
- 关键实验
- 评估了十个前沿模型,使用 202 个基准条目,涵盖六个法律领域和 24 个美国司法管辖区
- 主要贡献
- 提供首个专门评估大模型对不完整法律咨询查询识别能力的基准
- 意义与局限
- 填补了法律 AI 评估的空白,有助于提升模型处理真实用户咨询的能力,但仅限于美国法律环境