StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows· StartupBench: 市场验证的端到端工作流 AI 代理基准测试
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-or
通过市场验证的 AI 创业产品来评估 AI 代理在实际端到端工作流中的表现
- 核心方法
- 通过分析已经被市场验证的 AI 创业产品的实际工作流及用户需求,构建了一个全新的端到端 AI 代理基准测试 StartupBench
- 适合谁读
- 适合研究者和工程师阅读,特别是那些关注 AI 代理实际应用效果的人
- 要解决的问题
- 现有的 AI 代理基准测试多半基于研究者的假设,未曾针对实际用户需求进行评估
- 关键实验
- 未提供
- 主要贡献
- 提供了一个更加贴近实际应用场景的基准测试,帮助评估 AI 代理在不同专业领域中的实用能力
- 意义与局限
- 意义在于推进 AI 代理的实际应用,影响可能包括提高不同行业对 AI 代理的接受度,局限在于仅基于现有产品的工作流,可能忽略了新兴的需求