跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-05-15精选AI 评分73

PAGER:弥合点精确几何图形界面控制中的语义-执行鸿沟

AI 导读

研究针对需要点级精度的几何图形界面控制任务,揭示了现有视觉-语言模型存在的语义-执行鸿沟:通用模型动作类型准确率高但任务成功率极低。为此,我们构建了包含4,906个问题、超过22.4万次像素级动作的PAGE Bench基准,并提出了拓扑感知智能体PAGER。该智能体通过依赖结构规划与像素级执行分解任务,结合像素接地监督调优与精度对齐强化学习,将任务成功率提升至最强通用基线的4.1倍,步骤成功率从GUI专用智能体的不足9%提高到62%以上,实现了点精确GUI控制的新突破。

推荐理由

GUI agent一直绕着精确点击走,这篇直接硬碰硬,把成功率从6%拉到62%,做CAD自动化或工业软件的团队可以重点关注。

正文

Authors:Jingxuan Wei, Xi Bai, Shan Liu, Caijun Jia, Zheng Sun, Xinglong Xu, Siyuan Li, Linzhuang Sun, Bihui Yu, Conghui He, Cheng Tan

View PDF HTML (experimental)

Abstract:Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving region-tolerant paradigm, where many nearby pixels inside the same component remain valid. Precise geometric construction breaks this assumption: actions must land on points in continuous canvas space rather than tolerant regions. Because geometric primitives carry ontological dependencies, a local coordinate error can induce cascading topological failures that distort downstream objects and invalidate the final construction. We identify this regime as precision-sensitive GUI tasks, requiring point-level accuracy, geometry-aware verification, and robustness to dependency-driven error propagation. To benchmark it, we introduce PAGE Bench, with 4,906 problems and over 224K process-supervised, pixel-level GUI actions. We further propose PAGER, a topology-aware agent that decomposes construction into dependency-structured planning and pixel-level execution. Pixel-grounded supervised tuning establishes executable action grammar, while precision-aligned reinforcement learning mitigates rollout-induced exposure bias through state-conditioned geometric feedback. Experiments reveal a pronounced Semantic-Execution Gap: general multimodal models can exceed 88% action type accuracy yet remain below 6% task success. PAGER closes this gap, delivering 4.1x higher task success than the strongest evaluated general baseline and raising step success rate from below 9% for GUI-specialized agents to over 62%, establishing a new state of the art for point-precise GUI control.
Comments: 27 pages, 11 figures, 3 tables
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.15963 [cs.AI]
  (or arXiv:2605.15963v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.15963

arXiv-issued DOI via DataCite

Submission history

From: Cheng Tan [view email]
[v1] Fri, 15 May 2026 13:55:05 UTC (2,147 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org