ai.hackcv
论文精选 65arXiv

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance· BioSecBench-Surveillance: AI病原体基因组监控基准

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 pe

AI 解读论文

提出AI病原体基因组监控基准,测试AI从原始测序数据推断正确分析流程的能力。

核心方法
构建了包含100个评估任务的基准,每个任务提供与人类分析师相同的原始数据和监控背景,评估AI推断出的分析流程的准确性。
适合谁读
研究者、工程师
要解决的问题
病原体基因组监控在数据分析上存在瓶颈,需要评估AI在该领域的表现。
关键实验
在3,962次评估尝试中,测试了16个模型-框架组合,最强配置仅通过了约50%的任务。
主要贡献
提供了一个可验证的基准测试工具,覆盖了从分类学到基因工程检测等多个任务类别,能够评估不同模型的表现。
意义与局限
为AI在病原体基因组监控领域的应用和发展提供了重要评估工具,但目前AI模型的表现仍需进一步提升。
领域:cs.AI作者:Harmon Bhasin、Kevin Flyangolts、Dianzhuo Wang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考