X:Rohan Paul (@rohanpaul_ai)· X:Rohan Paul (@rohanpaul_ai)·· 4 天前AI 评分50
MerchantBench 评测显示八款 LLM 长周期任务表现远逊人类,最好模型仅达人类 27.3%
AI 导读
MerchantBench 用 365 天电商模拟环境评测八款 LLM 智能体在长期连贯性上的表现,包括 GPT-5.6 Sol 和 Claude Opus 4.8。
正文
This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.
The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.
A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
来源:X:Rohan Paul (@rohanpaul_ai) · x.com