跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分38

CUA-SWE:当 Computer-Use Agent 遇上可视化软件工程

AI 导读

研究者推出 CUA-SWE,一个面向"计算机使用+软件工程"的 benchmark、环境与评测流水线,覆盖四个软件工程领域,要求智能体在同一任务内改代码与配置、执行命令、操作运行中的软件并查看视觉反馈。每个任务配有确定性的专属测试,用于验证改动后的软件是否满足需求并保持既定行为。该工作评估前沿智能体如何结合源码级执行与应用截图、图形交互,产出经过验证的软件变更。

正文

View PDF HTML (experimental)

Abstract:Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.
Comments: 66 pages
Subjects: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
Cite as: arXiv:2609.32600 [cs.AI]
  (or arXiv:2609.32600v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.32600

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Prince Zizhuang Wang [view email]
[v1] Sat, 26 Sep 2026 13:18:00 UTC (5,009 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org