ai.hackcv
论文精选 65arXiv

Pretraining Data Can Be Poisoned through Computational Propaganda· 预训练数据可被计算性宣传中毒

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.

AI 解读论文

预训练数据可被计算性宣传中毒,提出 HalfLife 分析方法。

核心方法
利用公开讨论接口作为传播恶意内容的途径,并引入 HalfLife 方法来估计这些内容在网页爬取和数据整理后的保留率。
适合谁读
研究者
要解决的问题
预训练数据中存在的计算性宣传攻击可能导致语言模型产生有害行为。
关键实验
使用 HalfLife 分析方法评估通过公开讨论接口注入的恶意内容在预训练数据中的可行性。
主要贡献
证明了预训练数据在更大规模和更异质的设置下可被中毒,并提供了检测方法。
意义与局限
强调了评估预训练数据中毒的重要性和可能性,指出了第三方网页内容作为攻击向量的风险,但方法的普适性和攻击的实际效果仍需进一步研究。
领域:cs.AI作者:Victoria Graf、Hannaneh Hajishirzi、Noah A. Smith
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考