ai.hackcv
论文精选 85arXiv

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications· 研究想法实施差距评估基准

A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. Idea

领域:cs.CL作者:Yiling Ma、Yilun Zhao、Sihong Wu
相关推荐
论文精选 82

Do speech foundation models really learn words?· 语音基础模型真的学会单词了吗?

Self-supervised speech foundation models are now used in a wide array of downstream applic…

领域:cs.CL作者:Robin Huo、Ewan Dunbar
📎 arXiv🕒 09-10 00:47🔗 arxiv.org
论文精选 75

ReCite: Agentic Reasoning for Faithful Citation· ReCite: 精确引用的代理推理

Accurate citations are the foundation of academic writing, tracing intellectual origins an…

领域:cs.CL作者:Yuyang Huang、Bobo Li、Jiajia Song
📎 arXiv🕒 09-09 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考