InfoOps Bench: A live information operations safety benchmark· InfoOps Bench:动态信息操作安全基准
In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for state-backed information operations. We draw on over 2,100 information operations from a live monitoring pipeline which tracks Russian, Chinese and Iranian state-backed information assets. Alongside this paper, we release a companion website that tracks the most prominent claims spread by state-backed media outlets, updated weekly, available from: pattrn.ai/research/infoopsbench. The dynamic nature of the benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations. Integrity scores, defined as the percentage
提出动态基准评估语言模型在信息操作中的可靠性
- 核心方法
- 基于动态监控管道的2100多个信息操作案例,测试17种模型的4种提示框架
- 适合谁读
- 研究者、政策制定者、AI开发工程师
- 要解决的问题
- 防止国家支持的媒体利用语言模型进行信息操作
- 关键实验
- 测试了17个来自8个供应商的语言模型,每周更新最突出的国家媒体主张
- 主要贡献
- 提供了一个动态更新的基准测试工具和网站,评估语言模型的完整性
- 意义与局限
- 提高了对AI在政治宣传中潜在风险的认识,对模型完整性评估具有重要影响,但可能存在监控范围有限的问题