gemma-4-12B-it-qat-GGUF (Unsloth)
Google Gemma 4 12B QAT GGUF,本地多模态推理
社区实测
社区普遍认为这是目前性价比极高的本地编程模型,QAT 让低比特量化后质量损失极小,在消费级显卡上能跑到 70+ t/s。Q5_K_XL 被多位用户推荐为日常主力,兼顾速度与代码质量。部署门槛低,2-bit 仅 4.66 GB,16GB 显存即可运行。
- 大幅降低显存需求,可在 16GB 消费级 GPU 上本地运行 12B 模型
- 2-bit 量化仅占 4.66 GB 磁盘空间,部署门槛极低
- Q4_K_M 量化在 RTX 4080 Super 上可达 72.3 t/s 生成速度
- QAT 训练使 4-bit 格式内存占用降低约 72% 同时保持接近原始性能
- 开箱即用,设置缓存和上下文长度后即可接入工作流
- Q5_K_XL 量化下多数编程任务可一次通过,减少反复修改
- 支持 llama.cpp、Ollama、Transformers 等主流本地推理工具
- MTP(多 token 预测)在 llama.cpp 中的支持仍在开发中,尚未就绪
- UD 量化格式比标准格式大约 10%,速度慢 1-5%
- Unsloth 版本本质上是转换 Google 官方 QAT,并非独立改进
- Q8_0 量化速度骤降至 25.2 t/s,相比 Q4_K_M 的 72.3 t/s 差距明显
- 从 Q4 升到 Q5 后速度从 61 t/s 降至 50 t/s,需权衡质量与速度
2-bit Gemma 4 12B GGUF is amazing! (4.66 GB on disk) : r/unslothGemma 4 QAT GGUFs from Unsloth : r/LocalLLaMA - RedditGemma 4 QAT | Unsloth DocumentationGemma 4 12B is my new main squeeze : r/LocalLLaMA - RedditGemma 4 with quantization-aware training - Google BlogGemma 4 UD performance? : r/unsloth - RedditGoogle Official QAT vs. Unsloth Version: Which Gemma 4 12B ... - noteGemma 4 12B: A unified, encoder-free multimodal modelUnsloth Gemma 4 QAT MTP assistant models now available - Reddit
截至 2026-06-21
快速上手
llama.cpp:./llama-cli -m gemma-4-12B-it-qat-Q4_K_M.gguf