DeepSeek-V4-Flash GGUF (DeepSeek)
ds4引擎 DeepSeek-V4-Flash 量化,Mac 本地运行
仓库标识antirez/deepseek-v4-gguf
许可 可商用任务 文本生成最近核对 2026-08-05
- 运行时
- OLLAMA
- 下载
- 811,544
快速上手示例
ollama run hf.co/antirez/deepseek-v4-gguf 依赖版本和硬件参数请以源仓库说明为准。
适合与不适合
独立证据与社区反馈
社区普遍认为 DeepSeek-V4-Flash 0731 性价比极高,基准测试已逼近甚至超越部分旗舰模型(MiMo 2.5 Pro、GLM-5.2 级别),API 推理速度远超同规模开源模型。但本地部署在 llama.cpp 下推理速度偏慢(<15 tok/s),且实际使用体验不如 benchmark 分数稳定,独立评测覆盖仍不足。
- 复杂推理任务(ML 分类器/图)可达 MiMo 2.5 Pro 水平,200K 上下文内表现稳定
- 价格/性能比优于 Luna,兼具旗舰级质量与低成本
- 284B 总参数仅 13B 激活,MoE 架构可在消费级/爱好者硬件上以可接受速率运行
- API 输出速度 115.9 tok/s,远超同规模开源模型中位数(65.9 tok/s)
- 支持 1M token 上下文窗口
- 输出 UI 品味接近 Claude Opus 水平
- Terminal Bench 2.1 得分 82.7,DeepSWE 从 7.3 跃升至 54.4,大幅超越 V4-Pro-Preview
- 运行 30 分钟资源消耗几乎不增加(<1%),效率极高
- 通过 antirez DS4 DwarfStar 推理引擎可实现高效本地部署
- 推理效率相比 V3.2 仅需 27% FLOPs 和 10% KV cache
- 可在多模型工作流中作为研究/worker 模型稳定运行
- llama.cpp 下推理速度低于 15 tok/s,远不如专用引擎
- 双 RTX 3060 + 96GB RAM 配置下 IQ2_M 量化仅约 3.5 tok/s
- 0731 版本相比原版 V4 Flash 出现约 10 倍成本增加的问题
- benchmark 表现强但实际使用稳定性不足,存在「benchmark maxed」现象
- 社区量化版本种类有限,Bartowski/unsloth 等主流量化作者尚未发布多种量化
- 独立评测覆盖不足,尚无足够非生成式评测以支撑公开排名
- 0731 版本无架构变更,此前 GGUF 的显存占用数据仍适用
- 社区实测r/LocalLLaMA on Reddit: Some deepseek-v4-flash 20260731 opinion review
- 社区实测r/LocalLLaMA on Reddit: DeepSeek-V4-Flash-0731 unsloth gguf on A100
- 社区实测r/LocalLLaMA on Reddit: New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna
- 社区实测r/LocalLLaMA on Reddit: DeepSeek-V4-Flash has been updated
- 社区实测r/LocalLLaMA on Reddit: DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP Head
- 社区实测r/LocalLLaMA on Reddit: Deepseek V4 Flash 2, 3 and 4 bits GGUFs
- 社区实测r/LocalLLaMA on Reddit: deepseek-ai/DeepSeek-V4-Flash-0731 on Huggingface
- 社区实测r/DeepSeek on Reddit: Deepseek v4 flash 0731 real experience
- 独立评测DeepSeek V4 Flash Review (2026) — Specs, Tests & Speed
- 独立评测DeepSeek V4 Flash Benchmarks & Pricing (August 2026) | BenchLM.ai
- 独立评测DeepSeek V4 Flash 0731 (max) - Intelligence, Performance & Price Analysis
- 独立评测DeepSeek Retrained V4-Flash Beats Its Flagship Pro on Nine Agent Benchmarks
- 独立评测DeepSeek V4 Flash 0423 - API Pricing & Benchmarks | OpenRouter
- 社区实测r/LocalLLaMA on Reddit: Deepseek V4 Flash 0731 benchmarking
- 社区实测r/LocalLLaMA on Reddit: DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks
- 转载/仓库deepseek-ai/DeepSeek-V4-Flash · Hugging Face
可信度HF 81万下载,413赞,由Redis作者antirez发布
完整规格
下载动量
30天下载 1919.1k → 1957.6k · likes +42
观测时间线
模型家族
- DeepSeek-V4-Flash
- 量化 DeepSeek-V4-Flash GGUF (DeepSeek)