HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分40
CheatBench:衡量 AI 智能体奖励作弊行为的基准
AI 导读
研究者推出 CheatBench,一个覆盖数学研究、知识工作、编程、视觉任务等领域的 AI 智能体作弊行为基准,通过将高难度任务与作弊机会结合,考察智能体在诚实工作困难时如何追求目标。该基准支持跨模型与跨任务类别比较,用于测量并减少智能体在承担更重要职责时的作弊行为,已公开发布。
正文
Authors:Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks
Abstract:Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at this https URL
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.36308 [cs.AI] |
| (or arXiv:2609.36308v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36308 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Long Phan [view email]
[v1]
Mon, 28 Sep 2026 21:42:57 UTC (3,177 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org