跳到正文
原文
X:Elvis Saravia (@omarsar0, DAIR.AI)· X:Elvis Saravia (@omarsar0, DAIR.AI)·· 4 天前AI 评分47

webAI 开源 3.6B 模型 TwIL-LM3-Pro,BIG-Bench Hard 得分 95.4

AI 导读

webAI 发布开源模型 TwIL-LM3-Pro,仅 3.6B 参数,可在普通电脑本地量化运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:在形式逻辑上微调、将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升的同时通用推理保持稳定。该模型在 SVAMP 上得 95%、MuSR 上得 64.1%,均为所比较小模型中的最高分。

正文

Small models are getting really good at reasoning.

It's exciting because SLMs can unlock so much at the harness layer.

TwIL-LM3-Pro from @thewebAI has 3.6B parameters and runs locally on everyday computers. It scores 95.4 on BIG-Bench Hard, well ahead of Qwen3-8B at 63.7.

I like their post-training recipe. They fine-tune on formal logic, merge the weights back toward the base model, and then run RL against a programmatic verifier. Logic scores go up, and general reasoning holds steady.

Great to see more of this work released as open source.

来源:X:Elvis Saravia (@omarsar0, DAIR.AI) · x.com