Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies· 重建:从预出版参考文献恢复研究思想的盲基准
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review wit
通过预出版参考文献恢复研究思想的模型评估
- 核心方法
- 构建了Reconstruction基准测试,采用严格的防泄漏协议,隐藏源论文和所有同期或未来文献,让模型提出假设并由独立的大语言模型评判
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 评估语言模型能否仅通过预出版参考文献准确恢复研究论文的核心思想
- 关键实验
- 在六个科学领域对643篇论文进行了评估,结果显示前沿模型的匹配率仅为3-15%
- 主要贡献
- 提出了一个新的基准测试方法来评估语言模型对研究思想的恢复能力
- 意义与局限
- 该研究揭示了当前语言模型在恢复研究思想方面存在显著局限,为未来研究提供方向