ai.hackcv
论文精选 65arXiv

InfoOps Bench: A live information operations safety benchmark· InfoOps Bench:动态信息操作安全基准

In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for state-backed information operations. We draw on over 2,100 information operations from a live monitoring pipeline which tracks Russian, Chinese and Iranian state-backed information assets. Alongside this paper, we release a companion website that tracks the most prominent claims spread by state-backed media outlets, updated weekly, available from: pattrn.ai/research/infoopsbench. The dynamic nature of the benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations. Integrity scores, defined as the percentage

AI 解读论文

提出动态基准评估语言模型在信息操作中的可靠性

核心方法
基于动态监控管道的2100多个信息操作案例,测试17种模型的4种提示框架
适合谁读
研究者、政策制定者、AI开发工程师
要解决的问题
防止国家支持的媒体利用语言模型进行信息操作
关键实验
测试了17个来自8个供应商的语言模型,每周更新最突出的国家媒体主张
主要贡献
提供了一个动态更新的基准测试工具和网站,评估语言模型的完整性
意义与局限
提高了对AI在政治宣传中潜在风险的认识,对模型完整性评估具有重要影响,但可能存在监控范围有限的问题
领域:cs.AI作者:Dorian Quelle、Lisa-Maria Neudert、Jonathan Bright
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考