跳到正文
原文
X:Rohan Paul (@rohanpaul_ai)· X:Rohan Paul (@rohanpaul_ai)·· 4 天前AI 评分50

MerchantBench 评测显示八款 LLM 长周期任务表现远逊人类,最好模型仅达人类 27.3%

AI 导读

MerchantBench 用 365 天电商模拟环境评测八款 LLM 智能体在长期连贯性上的表现,包括 GPT-5.6 Sol 和 Claude Opus 4.8。

正文

This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.

The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.

A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.

来源:X:Rohan Paul (@rohanpaul_ai) · x.com