LongCat-Flash-Lite-Sparse (美团)
69B MoE,仅3B激活,1M上下文,面向Agent与长文本任务
仓库标识meituan-longcat/LongCat-Flash-Lite-Sparse
~41GB 4-bit许可 可商用任务 文本生成最近核对 2026-08-05
- 格式
- 4-bit
- 文件
- ~36GB
- 运行内存
- ~41GB–54GB
- 运行时
- PYTHON3
- 下载
- 155
仅 safetensors · 无 pickle 加载风险
快速上手示例
python3 -m sglang.launch_server --model meituan-longcat/LongCat-Flash-Lite-Sparse --trust-remote-code --chunked-prefill-size 2048 --nsa-prefill-backend fa3 --kv-cache-dtype bfloat16 依赖版本和硬件参数请以源仓库说明为准。
适合与不适合
独立证据与社区反馈
在 24GB 显存约束下是纯 instruct 场景的较好选择,API 速度快但本地部署门槛高
- 24GB 显存内可运行 68.5B 参数的纯 instruct 模型
- API 服务翻译体验好(400 tokens/s)
- 在 Agent 和编码场景中优于同参数规模的 MoE 基线模型
- 本地部署困难,Hugging Face 上仅有 MLX 版本
- 社区实测r/LocalLLaMA on Reddit: LongCat-Flash-Lite 68.5B maybe a relatively good choice for a pure instruct model within the 24GB GPU VRAM constraint.
- 社区实测r/LocalLLaMA on Reddit: meituan-longcat/LongCat-Flash-Lite
可信度MIT 许可,SGLang 已合入部署 PR #32918,69B/3B 激活
完整规格
下载动量
30天下载 1.3k → 908 · likes +15