X:Unsloth (@UnslothAI)· X:Unsloth (@UnslothAI)·· 2026-09-02精选AI 评分66
Unsloth 让 Qwen3.8-Flash 本地推理提速 1.7 倍
AI 导读
Unsloth 通过 MTP(Multi-Token Prediction)让 Qwen3.8-Flash-Next 本地推理提速约 1.3 至 1.7 倍且精度不变,GGUF 在单张 RTX PRO 6000 上可达 170 tokens/s(基线 100 tokens/s)。
推荐理由
原文给出了 MTP 加速的具体倍数、显存门槛和各量化档位内存需求,读者可以据此判断自己的设备能否跑通。
正文
Qwen3.8-Flash 现在借助 MTP 在本地运行速度提升 1.7 倍!⚡️
GGUF 在 RTX PRO 6000 上可达 170 tokens/s。
MTP 让 Qwen3.8-Flash-Next 的推理速度提升约 1.3–1.7 倍,且精度不变。
GGUF:https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
指南:https://unsloth.ai/docs/models/qwen3.8-next
来源:X:Unsloth (@UnslothAI) · x.com