Qwen3.8-27B-DFlash2 (incoai)
Qwen3.8 草稿模型,用于 SGLang/vLLM 投机解码
仅 safetensors · 无 pickle 加载风险
快速上手示例
python -m sglang.launch_server --model-path Qwen/Qwen3.8-27B --speculative-algorithm DFLASH --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 --speculative-num-draft-tokens 8 依赖版本和硬件参数请以源仓库说明为准。