Gemma-4-E4B-it-QAT (Google)
Google Gemma 4 量化版,本地消费级硬件推理
社区实测
社区普遍认为 Gemma 4 E4B QAT 是同尺寸中最强的小模型之一,QAT 让 4-bit 量化后质量损失极小,仅需 6GB 显存即可跑 agentic 工作流。指令遵循能力超出基准预期,但复杂推理任务上仍不及 GPT-5.4 等大模型,且部分用户日常更偏好 Qwen 3.6。
- 仅需 6GB 内存即可运行 agentic 工作流(网页搜索、代码执行)
- QAT 在 4-bit 量化下保持模型质量
- 指令遵循能力优于基准测试所暗示的水平
- Apache 2.0 许可证支持商用
- 首发即支持 Ollama、LM Studio、llama.cpp、Unsloth 等主流工具
- 支持 256K 上下文窗口
- 原生函数调用支持
- 可在 6-8GB VRAM 的笔记本上运行
- 复杂任务上不如 GPT 5.4、kimi 2.5 等更大模型
- 部分用户在日常使用中更偏好 Qwen 3.6
- 选错 Gemma 4 变体会导致体验很差
- 实际运行速度可能低于基准测试数据
来源
Gemma 4 with quantization-aware training : r/LocalLLaMAGemma 4 E4B is amazing! The 4-bit GGUF can web-search, execute ...Gemma 4 E4B - Am I missing something? : r/LocalLLMHonestly, Gemma 4 feels way better than the benchmarks sayHave you tried the Gemma 4 series, out of curiosity?Google Gemma 4: A Technical OverviewPick the Wrong Gemma 4 and You'll Think It's BrokenGemma 3 Was Dead Last. Gemma 4 Is World-Class. Here's ...Gemma 4: Byte for byte, the most capable open modelsGoogle Gemma 4 Developer Guide: Benchmarks & Local ...
截至 2026-06-21
快速上手
llama.cpp -m gemma-4-E4B-it-QAT.gguf