跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-23精选AI 评分73

腾讯发布 WorkBuddy Bench:多领域编码智能体评测套件

AI 导读

腾讯推出 WorkBuddy Bench,一个覆盖 Code、Web、Office、Security 四个工作领域的编码智能体评测套件。每个任务均从真实 commit、PR 或业务场景逆向工程而来,改写为口语化角色扮演请求,从构造上抵抗数据污染。该基准在 CodeBuddy Code 和 Claude Code 上运行,所有任务目录、环境镜像、评分工具和参考方案均开源发布。

推荐理由

腾讯发布的这个多领域编码代理基准,把代码、网页、办公、安全四块放进一套评估体系,反污染构建方式值得关注,做agent评估的团队可以认真看下。

正文

Authors:Tencent WorkBuddy Bench Team: Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, Zhijian Shao, Yuchen Shi, Shuwen Zhang, Chaofan Qiu, Linjie Che, Xiaoxi Zhao, Feng Wu, Kai Zhang, Chaofan Zhu, Yubin Qi, Xiaoyun Liang, Peijie Dong, Yunhao Zhang, Yuanjie Zhu, Ling Jiang, Xianjun Zhang, Zhehang Chu, Anyuan Sang, Zhen Feng, Sen Nie, Shi Wu, Yuanzhen Xu, Xin Li, Ning Yang, Zhiqiang Dong, Hande Dong, Qiang Lin, Yi Liu, Yunsheng Wu, Ke Li, Xing Sun

View PDF HTML (experimental)

Abstract:We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.
Comments: 30 pages, 9 figures. Project page: this https URL ; code: this https URL ; dataset: this https URL
Subjects: Computation and Language (cs.CL); Software Engineering (cs.SE)
Cite as: arXiv:2607.20911 [cs.CL]
  (or arXiv:2607.20911v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.20911

arXiv-issued DOI via DataCite

Submission history

From: Zhiheng Lyu [view email]
[v1] Thu, 23 Jul 2026 04:34:06 UTC (27,379 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org