HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-04-14精选AI 评分70
现实世界威胁下的移动 GUI 智能体:我们准备好了吗?
AI 导读
研究人员推出了一款应用内容插桩框架及配套测试套件,包含122个动态任务和3000余个静态场景,用于评估移动 GUI 智能体在含第三方内容(如广告、用户生成内容)的真实环境中的表现。实验显示,现有开源及商业智能体均受显著影响,在动态和静态环境下的平均误导率分别为42.0%和36.1%,表明当前技术在大规模部署前仍需加强安全性验证。
推荐理由
这篇论文系统揭示了移动 GUI 智能体的安全盲区,第三方恶意内容只需平均 10 个 token 就能让智能体误操作,视觉模态反而更脆弱,安全部署还有很长的路。
正文
Abstract:Recent years have witnessed a rapid development of mobile GUI agents powered by large language models (LLMs), which can autonomously execute diverse device-control tasks based on natural language instructions. The increasing accuracy of these agents on standard benchmarks has raised expectations for large-scale real-world deployment, and there are already several commercial agents released and used by early adopters. However, are we really ready for GUI agents integrated into our daily devices as system building blocks? We argue that an important pre-deployment validation is missing to examine whether the agents can maintain their performance under real-world threats. Specifically, unlike existing common benchmarks that are based on simple static app contents (they have to do so to ensure environment consistency between different tests), real-world apps are filled with contents from untrustworthy third parties, such as advertisement emails, user-generated posts and medias, etc. ... To this end, we introduce a scalable app content instrumentation framework to enable flexible and targeted content modifications within existing applications. Leveraging this framework, we create a test suite comprising both a dynamic task execution environment and a static dataset of challenging GUI states. The dynamic environment encompasses 122 reproducible tasks, and the static dataset consists of over 3,000 scenarios constructed from commercial apps. We perform experiments on both open-source and commercial GUI agents. Our findings reveal that all examined agents can be significantly degraded due to third-party contents, with an average misleading rate of 42.0% and 36.1% in dynamic and static environments respectively. The framework and benchmark has been released at this https URL.
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2507.04227 [cs.CR] |
| (or arXiv:2507.04227v2 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2507.04227 arXiv-issued DOI via DataCite |
|
| Related DOI: | https://doi.org/10.1145/3745756.3809249
DOI(s) linking to related resources |
Submission history
From: Guohong Liu [view email]
[v1]
Sun, 6 Jul 2025 03:31:36 UTC (3,559 KB)
[v2]
Tue, 14 Apr 2026 14:47:25 UTC (7,024 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org