ai.hackcv
论文精选 65arXiv

ContextWeave: A Real-World Workflow Benchmark· ContextWeave:真实世界工作流基准

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components

AI 解读论文

提出了一种真实世界办公流程的长期记忆基准测试。

核心方法
构建了一个名为ContextWeave的基准,包含14名参与者的多个月隐私保护工作流,共1,005个可执行任务,通过多维度评估记忆对下游任务性能的影响。
适合谁读
研究者、工程师、产品开发者
要解决的问题
现有评估方法未能充分测试语言代理在长时间、状态保持的工作流中的记忆能力。
关键实验
涉及六种记忆组件的评估,测试了任务的相关性、连续性、可解决性和回忆误导的鲁棒性,具体实验数据及结果未详述。
主要贡献
提供了一个全新的、更接近真实办公环境的记忆评估基准,填补了现有评估方法的空白。
意义与局限
为语言代理的记忆评估提供了更为实用的标准,有助于推动该领域技术的发展,但可能受限于特定办公场景。
领域:cs.AI作者:Bo Wang、Yuqian Yao、Enxi Wang
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考