跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-08-18精选AI 评分72

StartupBench:面向市场验证端到端工作流的通用智能体基准测试

AI 导读

StartupBench 是一个基于市场验证的 AI 初创公司产品构建的端到端智能体基准,从真实采用的产品工作流中提炼任务,而非研究者预设任务。在统一智能体框架下,最强模型也仅能完成约 30% 的任务,复杂指令遵循和领域专业知识是主要失败来源。该基准揭示了当前通用智能体在真实用户任务上的能力边界。

推荐理由

不同于基于研究者自选任务的基准,StartupBench 从有真实用户的产品工作流提炼任务,给出的约 30% 完成率将团队评估焦点从榜单得分拉回到端到端交付。

正文

Authors:Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

View PDF HTML (experimental)

Abstract:Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2608.17800 [cs.AI]
  (or arXiv:2608.17800v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2608.17800

arXiv-issued DOI via DataCite

Submission history

From: Jingzhe Ding [view email]
[v1] Tue, 18 Aug 2026 14:01:32 UTC (5,347 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org