跳到正文
原文
Gary Marcus:The Road to AI We Can Trust(RSS)· Gary Marcus:The Road to AI We Can Trust(RSS)·· 2026-09-04精选AI 评分74

Gary Marcus 评 GPT-6 Astra:进步明显但鲁棒性与可监控性存疑

AI 导读

Gary Marcus 发文点评 GPT-6 Astra,称多项报告显示其为真正的进步,OpenAI 产品显式创建并操纵符号世界模型,令其近十年的主张获得印证。

推荐理由

作者结合自身近十年主张神经符号世界模型的立场,指出 Astra 的关键未知在鲁棒性与可监控性,判断有具体依据。

正文

X avatar for @arcprize

ARC Prize@arcprize

GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:

7:39 PM · Sep 3, 2026 · 293K Views

51 Replies · 194 Reposts · 1.81K Likes

Hot take on OpenAI GPT-6 Astra*1, with a challenge to Greg Brockman’s claims about it being AGI toward the end:

  • Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.

  • As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations.

    X avatar for @arcprize

    ARC Prize@arcprize

    Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: "L8: hub q2 (8↓). Lengths: 14=1…" - It mapped operations to exact controls: …

    7:39 PM · Sep 3, 2026 · 19K Views

    2 Replies · 5 Reposts · 153 Likes

  • What we don’t know is how robust that capability is. That is THE key question.

  • Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.

  • And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.

  • Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.

  • As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.

  • The new system appears to be less monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also more alignable, not sure why.

  • Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)

1

This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.

Share

No posts

来源:Gary Marcus:The Road to AI We Can Trust(RSS) · garymarcus.substack.com