HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分47
AutoGUIWorld:用图像生成器作为 GUI 智能体的视觉世界模型
AI 导读
AutoGUIWorld 是一个数据生成框架,用图像生成器的视觉先验加规划器的任务知识合成 GUI 交互轨迹,无需部署或运行真实软件环境。它从操作系统上下文、视觉外观和界面状态的结构化规格中采样初始场景,经动作 grounding 与转移级质量过滤,在 Ubuntu、Windows、macOS 和 Chrome 上产出 79,266 条带空间标注的 step 级训练样本。
正文
Authors:Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang, Beiduo Chen, Muxi Chen, Chenchen Zhao, Hexuan Deng, Haolin Yang, Geyuan Zhu, Sa Zhu, Jianhuan Zhuo, Qiuyong Xiao, Jianhao Ruan, Yiran Peng, Jiayi Zhang, Tian Ye, Xinlei Yu, Tianwen Jiang, Jihong Zhang, Yuyu Luo
Abstract:GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.01215 [cs.CV] |
| (or arXiv:2610.01215v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01215 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Cheng Yang [view email]
[v1]
Thu, 1 Oct 2026 07:22:53 UTC (46,478 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org