ContextWeave: A Real-World Workflow Benchmark· ContextWeave:真实世界工作流基准
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components
提出了一种真实世界办公流程的长期记忆基准测试。
- 核心方法
- 构建了一个名为ContextWeave的基准,包含14名参与者的多个月隐私保护工作流,共1,005个可执行任务,通过多维度评估记忆对下游任务性能的影响。
- 适合谁读
- 研究者、工程师、产品开发者
- 要解决的问题
- 现有评估方法未能充分测试语言代理在长时间、状态保持的工作流中的记忆能力。
- 关键实验
- 涉及六种记忆组件的评估,测试了任务的相关性、连续性、可解决性和回忆误导的鲁棒性,具体实验数据及结果未详述。
- 主要贡献
- 提供了一个全新的、更接近真实办公环境的记忆评估基准,填补了现有评估方法的空白。
- 意义与局限
- 为语言代理的记忆评估提供了更为实用的标准,有助于推动该领域技术的发展,但可能受限于特定办公场景。