ai.hackcv
论文精选 65arXiv

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents· 什么是好的代理数据?

LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object $(E,q,τ,v)$, comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor a

AI 解读论文

提出代理数据的两层框架与生成范式,提升 LLM 代理交互数据质量。

核心方法
研究者提出将代理数据表示为一个可分解的对象 $(E,q,τ,v)$,并根据主要锚点组织生成范式,其中 E 代表环境规范,q 代表任务信号,τ 代表交互实现,v 为可选验证器。
适合谁读
研究者 / 工程师
要解决的问题
现有的 LLM 代理数据生成方法在不同领域中组织和评估方式各异,导致难以看清通用生成机制,同时将候选构建与验证选择混淆。
关键实验
未提供具体的实验数据。
主要贡献
提供了代理数据生成的通用框架,有助于理解、评估和改进 LLM 代理数据的质量。
意义与局限
该框架使研究者和工程师能够更系统地研究代理数据生成问题,但可能需要更多的实证研究来验证其有效性。
领域:cs.AI作者:Xingshan Zeng、Zishan Xu、Boju Zhang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考