ai.hackcv
论文精选 65arXiv

It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention

Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently co-occur, and MAs can pose challenges for low-bit quantization. In this study, we analyze the factors underlying AS and MAs that emerge at the initial position regardless of the token occupying it. Our experiments suggest that self-concentration of attention, resulting from the causal mask, and the subsequent Value-non-mixing in attention outputs contribute to AS and MAs. These findings provide new empirical evidence on the internal dynamics of LLMs, offering insights that may inform future quantization strategies and advance our understanding of the internal mechanisms of attention layers.

领域:cs.CL作者:Raito Kiya、Satoki Ohashi、Kosuke Sato
相关推荐
论文精选 75

ReCite: Agentic Reasoning for Faithful Citation· ReCite: 精确引用的代理推理

Accurate citations are the foundation of academic writing, tracing intellectual origins an…

领域:cs.CL作者:Yuyang Huang、Bobo Li、Jiajia Song
📎 arXiv🕒 09-09 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考