HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分41
WorldAuditBench:用多模态智能体审计交互式 3D 世界
AI 导读
研究者推出 WorldAuditBench,一个面向 3D 世界审计的基准,包含 13 个基于 Unreal Engine 5 和 Three.js 构建的环境中的 213 个异常任务,覆盖五类异常。在固定探索预算下评测五款前沿模型,两种审计范式的成功率仅为 6.6% 至 42.3%,远低于人类的 83.4%。该基准用于研究多模态智能体如何在交互式 3D 环境中耦合动作与视觉推理。
正文
Abstract:As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and this http URL, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.40325 [cs.AI] |
| (or arXiv:2609.40325v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.40325 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ziyan Jiang [view email]
[v1]
Wed, 30 Sep 2026 17:55:29 UTC (47,272 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org