Gemma-4-31B-it-GGUF (LM Studio)
Google 31B 指令模型 GGUF 量化版,本地消费级硬件推理
社区实测
Gemma 4 31B 不是最聪明的本地模型,但因其综合可用性和离线能力成为许多人日常首选;一次性编码表现亮眼,工具调用场景则逊色不少;消费级硬件可跑但量化版本选择与 VRAM 策略直接影响体验。
- 本地离线推理,适合通勤、旅行、无网络环境
- 通过 GPU offloading 在消费级笔记本(如 RTX 5080 Laptop 16GB VRAM)上运行大参数量模型
- 多模态输入(文本+图像)
- 工具调用与推理能力
- 256K 长上下文窗口
- 通过推测解码(speculative decoding)加速推理
- 兼容多种推理框架:llama.cpp、Ollama、Unsloth Studio、Docker Model Runner 等
- lmstudio-community 的 GGUF 量化不在 KL 散度 Pareto 前沿(Q8_0 除外),品质不如 unsloth 或 bartowski 的量化
- 长文档和非拉丁脚本在低精度量化下退化最快,即使 Q8_0 也有明显 KL 散度
- 推测解码若草稿模型选择不当,速度反而比不用更慢(实测 7.31 t/s vs 57 t/s 基线)
- LM Studio 中 31B 版本的 thinking mode 开关可能不显示,需通过官网「use this model in LM Studio」按钮或手动修改聊天模板修复
- 编码任务需将 temperature 降至 0.3 或更低,默认 1.0 会产生大量错误
- 工具调用/自定义 harness 场景下表现不如一次性编码
- 密集模型对内存带宽要求高,不适合低带宽设备(如 NVIDIA Spark)
- 低精度量化下滑动窗口注意力与 KV cache 复用可能导致长上下文误差累积
Gemma 4 31B GGUF quants ranked by KL divergence ... - RedditSpeculative Decoding works great for Gemma 4 31B with E2B draft ...For anyone having issues with Gemma 4 31b in LM Studio ... - RedditGemma-4-26B-A4B-it-UD-Q4_K_M.gguf : IMHO worst model ever ...I ran Gemma 4 as a local model in Codex CLI - Hacker NewsGemma 4 31B GGUF Quality Benchmark: unsloth, bartowski ...Gemma 4 - LM StudioPractical Gemma 4 Benchmarking with LM Studio - DEV CommunityI Spent 3 Nights Testing Gemma 4 (MTP)Google's Gemma 4 isn't the smartest local LLM I've run ... - Facebooklmstudio-community/gemma-4-31B-it-GGUF - Hugging FaceSlow inference with 31b model Gemma 4? Optimizations?
截至 2026-06-21
快速上手
ollama run hf.co/lmstudio-community/gemma-4-31B-it-GGUF:Q4_K_M