ai.hackcv
论文精选 65arXiv

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing· BLOOM-WILT:大模型审计中的逻辑倾斜技术

Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of magnitude more interactions than any evaluation can simulate. Automated auditors make testing cheap to scale and flexible enough to cover almost any specified behaviour, yet their lack of optimisation pressure makes them sample-inefficient. To address this shortcoming, we introduce BLOOM-WILT, a full auditing pipeline that elicits natural multi-turn instances of rare behaviours, without training cost or access beyond the target's next-token distribution. On the input side, WILT's auditor model revises its conversational strategy across rounds, learning from previous scored interactions. On the output side, WILT adaptively reweights the targ

AI 解读论文

提高大语言模型审计效率的技术,无需额外训练成本。

核心方法
BLOOM-WILT 通过在输入端调整对话策略并从之前的交互中学习,在输出端自适应地重新加权目标的下一个词分布,以高效地审计罕见行为。
适合谁读
研究者 / 工程师
要解决的问题
解决大语言模型在部署后难以通过常规测试发现罕见行为的问题。
关键实验
未提供
主要贡献
提出了一种新的审计方法 BLOOM-WILT,能够在不增加训练成本和无需访问模型内部结构的情况下,高效地审计大语言模型的罕见行为。
意义与局限
BLOOM-WILT 提高了模型审计的效率和效果,有助于更好地发现和理解模型的潜在问题,但可能需要更多实验验证其效果。
领域:cs.AI作者:Adrians Skapars、Edoardo Manino
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考