Gemma4-Gutenberg-31B-Heretic-mlx-8Bit (ailexleon)
MLX 8-bit 量化英文创意写作模型,适合故事小说生成
- Apple Silicon
- 已量化
仅 safetensors · 无 pickle 加载风险
社区实测
Gemma 4 31B 基准表现强劲,以 31B 参数规模匹敌大数十倍的模型,但社区部署反馈两极:推理速度明显慢于 Qwen 3.5,KV 缓存占用大导致同等显存下可用的上下文长度远不及竞品;Heretic 去审查版本在创意写作场景受认可,但 8-bit 量化质量可能不如 4-bit。Apache 2.0 许可被社区视为比跑分更重要的企业友好信号。
- Heretic 版本通过 ARA 方法移除了 Gemma 4 的安全对齐,实现去审查输出
- 社区反馈 Gemma 4 在创意写作方面表现更好,Heretic 版本进一步降低了拒绝率
- Apache 2.0 许可证消除了企业对 Google 使用政策单方面变更的依赖风险
- 31B 参数规模即可在 Arena AI Elo 上匹敌 600B 级模型,大幅降低了本地运行门槛
- AIME 数学基准从 Gemma 3 的约 20-29% 跃升至 80% 以上,代际提升显著
- 已有面向 Apple Silicon 的 MLX 8-bit 量化版本,可在 Mac 上本地运行
- SWA 机制一定程度上缓解了 KV 缓存膨胀问题
- 8-bit 量化质量存疑:有测试显示 4-bit 得分与 16-bit 持平(21/23),而 8-bit 反而更低(20/23)
- 推理速度远慢于 Qwen 3.5:同硬件下仅 11 t/s,Qwen 3.5 可达 60 t/s
- KV 缓存占用巨大:完整 260K 上下文 fp16 需约 22GB 显存,远不如 Qwen 3.5 紧凑
- 同等显存下可用上下文远少于 Qwen 3.5(约 20K vs 190K)
- Google AI Studio 托管的 Gemma 4 版本被社区评价为体验很差
- 有用户认为 Gemma 4 在非编程任务上完全无用
TheCluster/Gemma-4-31B-Heretic-MLX-8bit - Hugging FaceGemma 4 31B — 4bit is all you need : r/LocalLLaMA - RedditGemma 4's defenses shredded by Heretic's new ARA method 90 ...anything better than gemma-4-26B-A4B-it-heretic-GGUF - RedditG4-Meromero-31B-Uncensored-Heretic Is Out Now, a Finetune of ...Why is noone here talking about Gemma4? It's been out for a while ...Gemma 4 is good : r/LocalLLaMA - RedditGemma 4 Benchmarks: How a 31B Model Competes with Giants 20× Its Size | Gemma4AllGemma 4's Real Breakthrough Isn't the Benchmarks Google just handed enterprises something worth…TxemAI/gemma-4-31B-uncensored-heretic-mlx-8bit · Hugging FaceGitHub - Incept5/gemma4-benchmark: MLX benchmark: Gemma 4 + Qwen 3.5 on Apple Silicon with TurboQuant KV cache · GitHub
截至 2026-07-05
快速上手
mlx_lm.generate --model ailexleon/Gemma4-Gutenberg-31B-Heretic-mlx-8Bit --prompt 'hello'