X:Unsloth (@UnslothAI)· X:Unsloth (@UnslothAI)·· 2026-08-27精选AI 评分72
Unsloth 发布 GLM-5.3-Flash GGUF 量化版本地运行方案
AI 导读
Unsloth 发布 GLM-5.3-Flash(ox-alpha)的 GGUF 量化版本,可在 128GB RAM 设备上以 3-bit 运行。该模型为 Z.ai 的 320B-A18B 开源多模态模型,MIT License 发布,原文称其在 DeepSWE、编码和智能体基准上媲美 Claude Opus 4.8。
推荐理由
原文给出了各量化档位的内存需求和精度保留数据,读者可据此判断在 128GB 设备上本地运行该模型的可行性。
正文
GLM-5.3-Flash 现在可以在本地运行了!✨
通过 Unsloth GGUF 在 128GB 内存上以 3-bit 量化运行。
GLM-5.3-Flash(ox-alpha)在 DeepSWE、编程和智能体基准测试上与 Claude Opus 4.8 不相上下。
指南:https://unsloth.ai/docs/models/glm-5.3-flash
GGUF:https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
来源:X:Unsloth (@UnslothAI) · x.com