ai.hackcv
论文精选 65arXiv

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text· 检测人类-大模型共创作文本中的大模型生成标志

The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs. This paper introduces a new method to address this urgent need. Our method operates at the token level, the natural unit of modern language models, and builds on existing token-level detection scores. The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure. Our method is simple to implement and does not require token

AI 解读论文

提出了一种平滑相邻词元得分以检测人类-大模型共创作文本中大模型生成部分的新方法。

核心方法
通过在词元级别平滑相邻词元的检测得分,并使用自适应Lepski规则选择带宽来适应局部作者结构,实现对大模型生成内容的细粒度检测。
适合谁读
研究者 / 工程师 / 产品
要解决的问题
现有的大模型生成文本检测方法主要集中在文档级别分类,无法精确定位混合作者文本中由大模型生成的具体部分。
关键实验
未提供
主要贡献
提供了一种简单易实现的词元级别检测方法,能够精确定位大模型生成的内容,适用于混合作者文本的细粒度分析。
意义与局限
该方法有助于提高人类-AI协作文本的透明度和可信度,但可能存在对不同类型大模型和文本风格的适应性局限。
领域:cs.AI作者:Yangjun Lu、Hongyi Zhou、Fabian Spill
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考