Can AI agents conduct open-ended AI research? Early evidence from two case studies· AI能进行开放性研究吗?两个案例的早期证据
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents compl
AI能否进行开放性研究的初步探索。
- 核心方法
- 通过给AI代理提供未发表的高质量论文中的核心开放性研究问题,并由原作者对其输出进行评分,提出了影子评估的第三种评估方式。
- 适合谁读
- 研究者,尤其是关注AI自动化研究开发领域的学者
- 要解决的问题
- 目前缺乏关于AI代理能否承担开放性AI研究任务的有力证据。
- 关键实验
- 对两个未发表的NeurIPS 2026论文进行了影子评估实验,给予AI代理六天时间及数万美元的计算资源。
- 主要贡献
- 提出了评估AI进行开放性研究能力的新方法——影子评估,为AI自动化研究开发提供了新的视角。
- 意义与局限
- 为评估AI在进行复杂科学研究方面的潜力提供了初步证据,但实验规模较小,计算成本高昂,存在一定的局限性。