ai.hackcv
论文精选 60arXiv

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models· 分子属性预测中的大模型文字级检索

Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged $89\%$ more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cas

AI 解读论文

大模型在分子属性预测中普遍存在文字级数值检索现象

核心方法
作者通过审计22个前沿模型在12个回归基准上的表现,分析了模型在不同推理水平下进行文字级数值检索的情况
适合谁读
研究者、工程师
要解决的问题
评估大语言模型在分子属性预测任务中的准确性时,难以区分模型是真正预测还是检索已发表的数值
关键实验
在12个回归基准上对22个模型进行审计,实验包含两个推理水平,展示了不同水平下检索频率的变化
主要贡献
揭示了大模型在某些分子属性预测数据集上存在显著的文字级检索现象,并提出了一种中断检索的方法
意义与局限
研究发现有助于理解大模型在特定任务上的行为机制,但限于现有数据集和实验设置,结论可能需要进一步验证
领域:cs.AI作者:Matthias Busch、Marius Tacke、Sviatlana V. Lamaka
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考