Qwen3.6-35B-A3B-GGUF (Unsloth)
支持MTP加速的3B激活参数量化视觉语言模型,面向本地编程代理
社区实测
社区普遍认为这是 Qwen3.5 的明显升级,工具调用能力在低比特量化下依然出色,Unsloth 的 Dynamic 量化质量在同类 GGUFs 中排名靠前,适合消费级硬件本地部署;但在纯 CPU 场景下推理速度可能比部分其他来源的 GGUF 慢约 30%。
- 2-bit 量化即可完成 30+ 次工具调用、搜索 20 个站点并执行 Python 代码,仅需 13GB RAM
- 1-bit GGUF 的工具调用表现依然很好
- SVG 绘图质量在特定场景下可媲美甚至超过 Claude Opus 4.7
- 开放式探索搜索任务相比 Qwen3.5 有明显提升
- Unsloth Dynamic 量化在 KL 散度基准上达到 SOTA(99.9%),覆盖编码、聊天、工具调用、科学、非拉丁文字等场景
- 支持广泛的推理框架(llama.cpp、vLLM、SGLang、LM Studio、Docker Model Runner 等)
- 可在 RTX 4060 8GB 显存上运行并获得不错的评估结果
- 纯 CPU 推理时 Unsloth GGUF 比同等大小的其他来源 GGUF 慢约 30%,后续响应处理时间也更长
- 量化模型中存在 ssm_conv1d 张量漂移问题
- SVG 细节仍有瑕疵(如火烈鸟坐在轮胎上而非车座上)
- 不同用户的体验差异较大
2-bit Qwen3.6-35B-A3B GGUF is amazing! Made 30+ successful tool calls : r/unslothI've been running this on my laptop with the Unsloth 20.9GB GGUF in LM Studio: h... | Hacker NewsQwen3.6-35B-A3B GGUF Performance Benchmarks. : r/unslothQwen3.6-35B-A3B GGUF from Unsloth is quite a bit slower? : r/LocalLLaMAunsloth/Qwen3.6-35B-A3B-GGUF - Hugging FaceQwen 3.6 35B A3B GGUF Quality Benchmark: unsloth, bartowski, lmstudio-community, ggml-org, mudler, AesSedai comparedQwen3.6-35B-A3B-Uncensored-Wasserstein-GGUF : r/LocalLLaMAGetting Crazy Eval using Unsloth Qwen3.6 35B A3B on a 4060 with ...Qwen3.6-35-A3B outperforms on 21 of 22 model sizes in GGUF ...Qwen3.6 - How to Run Locally | Unsloth DocumentationQwen3.6-35B-A3B on my laptop drew me a better pelican than ...Qwen3.5 GGUF Benchmarks | Unsloth Documentation
截至 2026-06-21
快速上手
llama-server -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL --spec-type draft-mtp --spec-draft-n-max 6