ai.hackcv
论文精选 83arXiv

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier· IBIB: 企业AI系统评估新协议

Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document,

领域:cs.CL作者:Blake Stenstrom、Charangan Vasantharajan、Brian Sathianathan
相关推荐
论文精选 82

Do speech foundation models really learn words?· 语音基础模型真的学会单词了吗?

Self-supervised speech foundation models are now used in a wide array of downstream applic…

领域:cs.CL作者:Robin Huo、Ewan Dunbar
📎 arXiv🕒 09-10 00:47🔗 arxiv.org
论文精选 75

ReCite: Agentic Reasoning for Faithful Citation· ReCite: 精确引用的代理推理

Accurate citations are the foundation of academic writing, tracing intellectual origins an…

领域:cs.CL作者:Yuyang Huang、Bobo Li、Jiajia Song
📎 arXiv🕒 09-09 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考