跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分47

Argo-Bench:面向企业级工作流的数据智能体评测基准

AI 导读

Argo-Bench 是一个包含 210 个数据科学与分析任务的企业级数据智能体评测框架,模拟纽约市外卖平台 2024 年 8100 万订单,导出为 235 张表、75 亿行的 ERP 仓库。任务要求智能体在仓库中重建事实并执行封禁欺诈账号、分配骑手激励预算等操作,由模拟器按后果评分。14 个前沿与开源权重模型中最强者仅在 34.8% 的任务上得分 95 以上,平均 59.5 分。

正文

View PDF HTML (experimental)

Abstract:Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
Comments: 41 pages, 4 figures, 18 tables. Code: this https URL. Data: this https URL. Website: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)
Cite as: arXiv:2610.02122 [cs.CL]
  (or arXiv:2610.02122v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02122

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Arman Raayatsanati [view email]
[v1] Thu, 1 Oct 2026 17:35:44 UTC (243 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org