跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分44

TraceDance:从真实部署轨迹自动构建智能体行为基准的自动化系统

AI 导读

TraceDance 是一个从真实部署轨迹自动构建智能体行为基准的系统,通过 Anchor-and-Confirm 与 Anchor Synthesis Loop 针对用户指定的不良行为生成测试。

正文

Authors:Dehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang, Ziyi Chen, Yan Zhang, Qinbo Bai, Mengyuan Chao, Jing Ning, Qiyue Hua, Huiyi Chen, Hanrong Zhang, Henry Peng Zou, Jie Yang, Wei Xu, Philip S. Yu

View PDF HTML (experimental)

Abstract:An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
Comments: 34 pages, 7 figures. Project website: this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.33295 [cs.AI]
  (or arXiv:2609.33295v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.33295

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dehai Min [view email]
[v1] Sun, 27 Sep 2026 06:58:39 UTC (723 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org