跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分40

GPT-6 Astra 机器人智能体:成功率提升 14%,Token 用量减少 65%

AI 导读

PyRUA-Lean 框架让 GPT-6 Astra 驱动的机器人智能体在同等 LLM 调用预算下,将 LIBERO-PRO、RoboTwin 2.0 和 RoboCasa365 共 700 个仿真任务的整体成功率从 63.1% 提升至 71.7%。

正文

Authors:Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou

View PDF HTML (experimental)

Abstract:Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.01939 [cs.CV]
  (or arXiv:2610.01939v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.01939

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruiyang Si [view email]
[v1] Thu, 1 Oct 2026 16:07:49 UTC (2,932 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org