ai.hackcv
论文精选 65arXiv

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding· OmegaUse-OfficeVal: 基于经济基础的办公任务大模型评估基准

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inf

AI 解读论文

基于经济基础评估大模型在办公套件长期任务中的表现

核心方法
引入OmegaUse-OfficeVal,包含100个任务,每个任务配有两组经济信号:人工劳动时间和任务价格代理
适合谁读
研究者、工程师、产品管理人员
要解决的问题
缺乏评估大模型在合理成本下执行办公套件长期任务的基准
关键实验
未提供
主要贡献
提供了一个评估大模型在办公任务中的经济效率和性能的新工具
意义与局限
推动大模型在实际办公场景中的应用,促进模型的商业化和效率评估;可能受限于任务复杂度和经济信号的准确性
领域:cs.AI作者:Jingbo Zhou、Yusai Zhao、Qi Bao
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考