指南 / 文本生成模型 / Apple Silicon 运行时
标注支持 MLX 运行时的文本生成模型
共 58 个
收发布方标注支持 MLX 运行时的模型(取 key_info.frameworks)。这只表示可以在 Apple Silicon 上用 MLX 跑,不代表仓库里有现成的 MLX 权重;需要现成权重请看「MLX 权重」。
Qwen3.5-9B-Atlassian-Q4-mlx (LeanZero)
LoRA · 基于 Qwen3.5-9B
面向 Atlassian Forge 的 6-bit MLX 模型
- 参数量
- 9B
- 商用
- 可商用
当前运行配置
Mihai-LeanZero/Qwen3.5-9B-Atlassian-Q4-mlx- 显存
- ~5.4GB 4-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model Mihai-LeanZero/Qwen3.5-9B-Atlassian-Q4-mlx --prompt "Which Forge module adds a panel to the Jira issue view?"- 许可证
- Apache-2.0
- 上下文
- 32k
- 国内可达
- 需代理
- 其它形态
- MLX·mihai-leanzero · MLX·mihai-leanzero
当前运行配置
Mihai-LeanZero/Qwen3.8-27B-Atlassian-Q8-mlx- 显存
- ~30GB 8-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model Mihai-LeanZero/Qwen3.8-27B-Atlassian-Q8-mlx- 许可证
- Apache-2.0
- 上下文
- 128k
- 国内可达
- 需代理
- 其它形态
- MLX·mihai-leanzero · MLX·mihai-leanzero
Qwen3.6-35B-A3B-VQ-3.4bpw (TheDrainFlorist)
量化自 Qwen3.6-35B-A3B
Apple 芯片 13.8GiB VQ 量化,24GB 可跑
- 参数量
- 35B-A3B
- 商用
- 可商用
当前运行配置
TheDrainFlorist/Qwen3.6-35B-A3B-VQ-3.4bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.6-35B-A3B-VQ-3.4bpw --prompt 'Explain the difference between a mutex and a semaphore.' --max-tokens 512- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- apache-2.0
- 国内可达
- 需代理
Qwen3.6-35B-A3B-VQ-3.8bpw (TheDrainFlorist)
量化自 Qwen3.6-35B-A3B
Apple Silicon VQ量化 Qwen3.6-35B-A3B
- 参数量
- 35B-A3B
- 商用
- 可商用
当前运行配置
TheDrainFlorist/Qwen3.6-35B-A3B-VQ-3.8bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.6-35B-A3B-VQ-3.8bpw- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 国内可达
- 需代理
Qwen3.6-35B-A3B-VQ-4.6bpw (TheDrainFlorist)
量化自 Qwen3.6-35B-A3B
Mac 可跑的 Qwen3.6-35B VQ 量化
- 参数量
- 35B-A3B
- 商用
- 可商用
当前运行配置
TheDrainFlorist/Qwen3.6-35B-A3B-VQ-4.6bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.6-35B-A3B-VQ-4.6bpw --prompt "Hello"- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- apache-2.0
- 国内可达
- 需代理
Qwen3.6-35B-A3B-VQ-5.4bpw (TheDrainFlorist)
量化自 Qwen3.6-35B-A3B
MLX 向量量化,Apple Silicon 本地推理
- 参数量
- 35B-A3B
- 商用
- 可商用
当前运行配置
TheDrainFlorist/Qwen3.6-35B-A3B-VQ-5.4bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.6-35B-A3B-VQ-5.4bpw --prompt "Hello" --max-tokens 32- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 国内可达
- 需代理
当前运行配置
TheDrainFlorist/Qwen3.8-27B-VQ-3.9bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.8-27B-VQ-3.9bpw --prompt "Explain vector quantization briefly." --max-tokens 512- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 国内可达
- 需代理
当前运行配置
TheDrainFlorist/Qwen3.8-27B-VQ-4.5bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.8-27B-VQ-4.5bpw --prompt "Explain vector quantization briefly." --max-tokens 512- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- apache-2.0
- 国内可达
- 需代理
当前运行配置
TheDrainFlorist/Qwen3.8-27B-VQ-4.8bpw- 显存
- 暂无估算
查看详情查看运行方式
python -m mlx_lm generate --model TheDrainFlorist/Qwen3.8-27B-VQ-4.8bpw --prompt "Explain vector quantization briefly." --max-tokens 512- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 国内可达
- 需代理
当前运行配置
XHToken/Spark-X2.5-1.7B- 显存
- 暂无估算
查看详情查看运行方式
vllm serve XHToken/Spark-X2.5-1.7B --trust-remote-code- 为何暂无估算
- 权重格式未知,无法估算显存
- 许可证
- apache-2.0
- 上下文
- 1M
- 国内可达
- 需代理
当前运行配置
Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed- 显存
- ~16GB 4-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed- 许可证
- Apache-2.0
- 上下文
- 262k
- 国内可达
- 需代理
当前运行配置
Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality- 显存
- ~30GB 8-bit(估算)
查看详情查看运行方式
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality- 许可证
- Apache-2.0
- 上下文
- 262k
- 国内可达
- 需代理
Qwen3.8-Flash-Next Optimized Speed (MTPLX)
量化自 Qwen3.8-Flash-Next
Mac 上 MTPLX 4-bit 量化,MTP 投机加速
- 参数量
- 125B-A6B
- 商用
- 需自查
当前运行配置
Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed- 显存
- ~75GB 4-bit(估算)
查看详情查看运行方式
mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed- 许可证
- qwen-community-1.0
- 上下文
- 262k
- 国内可达
- 需代理
当前运行配置
Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed- 显存
- ~16GB 4-bit(估算)
查看详情查看运行方式
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed- 许可证
- apache-2.0
- 上下文
- 262k
- 国内可达
- 需代理
Dagger-Qwen3.6-27B (peculiar-ragdoll)
量化自 ThinkingCap-Qwen3.6-27B
Mac 用 Qwen3.6-27B MLX 量化,低冗余输出
- 参数量
- 27B
- 商用
- 可商用
当前运行配置
peculiar-ragdoll/Dagger-Qwen3.6-27B-MLX- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model peculiar-ragdoll/Dagger-Qwen3.6-27B-MLX- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- Apache-2.0
- 上下文
- 256k
- 国内可达
- 需代理
当前运行配置
OBLITERATUS/Qwen3.8-27B-OBLITERATED- 显存
- 暂无估算
查看详情查看运行方式
ollama run hf.co/OBLITERATUS/Qwen3.8-27B-OBLITERATED- 为何暂无估算
- 权重格式未知,无法估算显存
- 许可证
- Apache-2.0
- 国内可达
- 需代理
当前运行配置
HamoAI/hamo-score-0.6b- 显存
- 暂无估算
查看详情查看运行方式
ollama run hf.co/HamoAI/hamo-score-0.6b- 为何暂无估算
- 权重格式未知,无法估算显存
- 许可证
- HAMO-RAIL-S 1.0
- 国内可达
- 需代理
当前运行配置
lmstudio-community/LFM2.5-2.6B-MLX-5bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model lmstudio-community/LFM2.5-2.6B-MLX-5bit- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- LFM Open License v1.0
- 国内可达
- 需代理
当前运行配置
lmstudio-community/LFM2.5-2.6B-MLX-6bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model lmstudio-community/LFM2.5-2.6B-MLX-6bit- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- LFM-1.0
- 国内可达
- 需代理
当前运行配置
LiquidAI/LFM2.5-2.6B- 显存
- 暂无估算
查看详情查看运行方式
llama-server -hf LiquidAI/LFM2.5-2.6B-GGUF- 为何暂无估算
- 权重格式未知,无法估算显存
- 许可证
- LFM Open License v1.0
- 上下文
- 128k
- 国内可达
- 需代理
- 其它形态
- 原生 · GGUF·liquidai · GGUF·lmstudio · 4bit · 8bit
当前运行配置
AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-6bit-MTP- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-6bit-MTP- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 上下文
- 262k
- 国内可达
- 需代理
当前运行配置
lmstudio-community/Ornith-1.0-9B-MLX-5bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model lmstudio-community/Ornith-1.0-9B-MLX-5bit- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- MIT
- 国内可达
- 需代理
当前运行配置
lmstudio-community/Ornith-1.0-35B-MLX-6bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model lmstudio-community/Ornith-1.0-35B-MLX-6bit- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- MIT
- 国内可达
- 需代理
当前运行配置
lmstudio-community/Ornith-1.0-9B-MLX-6bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model lmstudio-community/Ornith-1.0-9B-MLX-6bit- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- MIT
- 国内可达
- 需代理
当前运行配置
lmstudio-community/Ornith-1.0-35B-MLX-5bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model lmstudio-community/Ornith-1.0-35B-MLX-5bit- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- mit
- 国内可达
- 需代理
当前运行配置
Goekdeniz-Guelmez/JOSIE-2-4B-Preview- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model Goekdeniz-Guelmez/JOSIE-2-4B-Preview- 为何暂无估算
- 权重格式未知,无法估算显存
- 许可证
- MIT
- 国内可达
- 需代理
当前运行配置
FINAL-Bench/POCKET-KR-MLX- 显存
- ~12GB 2-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model FINAL-Bench/POCKET-KR-MLX- 许可证
- Apache-2.0
- 国内可达
- 需代理
当前运行配置
AtomicChat/ornith-35b-MLX-8bit- 显存
- ~39GB 8-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model AtomicChat/ornith-35b-MLX-8bit --prompt Hello- 许可证
- MIT
- 上下文
- 256K
- 国内可达
- 需代理
当前运行配置
AtomicChat/ornith-35b-MLX-6bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model AtomicChat/ornith-35b-MLX-6bit --prompt 'Hello'- 为何暂无估算
- 已有格式线索,但估算映射未支持,无法估算显存
- 许可证
- MIT
- 上下文
- 256K
- 国内可达
- 需代理
当前运行配置
majentik/gpt-oss-20b-TurboQuant-MLX-2bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model majentik/gpt-oss-20b-TurboQuant-MLX-2bit --prompt '你的提示'- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache 2.0
- 国内可达
- 需代理
当前运行配置
zak-raindog/gear-manual-qwen3-4b- 显存
- ~2.4GB 4-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model zak-raindog/gear-manual-qwen3-4b --prompt "如何使用EP-133?"- 许可证
- cc-by-nc-4.0
- 国内可达
- 需代理
当前运行配置
prism-ml/Bonsai-27B-mlx-1bit- 显存
- ~4.1GB 1-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model prism-ml/Bonsai-27B-mlx-1bit- 许可证
- apache-2.0
- 上下文
- 262K
- 国内可达
- 需代理
当前运行配置
mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-OptiQ-4bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-OptiQ-4bit- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- nvidia-open-model-license
- 国内可达
- 需代理
当前运行配置
mlx-community/Qwen3.5-35B-A3B-OptiQ-4bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model mlx-community/Qwen3.5-35B-A3B-OptiQ-4bit --prompt 'Explain quantum computing.' --max-tokens 200- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- apache-2.0
- 上下文
- 128k
- 国内可达
- 需代理
当前运行配置
prism-ml/Ternary-Bonsai-27B-mlx-2bit- 显存
- 暂无估算
查看详情查看运行方式
llama-server -hf prism-ml/Ternary-Bonsai-27B-gguf- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- apache-2.0
- 上下文
- 262k
- 国内可达
- 需代理
Qwen3.6-35B-A3B-VQ (aquaman164)
量化自 Qwen3.6-35B-A3B
MLX VQ量化35B MoE,适合Apple Silicon本地推理
- 参数量
- 35B (3B激活)
- 商用
- 可商用
当前运行配置
aquaman164/Qwen3.6-35B-A3B-MLX-VQ-2.6bpw- 显存
- 暂无估算
查看详情查看运行方式
huggingface-cli download aquaman164/Qwen3.6-35B-A3B-MLX-VQ-2.6bpw --local-dir qwen-vq && python qwen-vq/code/vq_serve.py --model qwen-vq- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 上下文
- 128k
- 国内可达
- 需代理
Ornith-1.0-35B-MLX (deepreinforce-ai)
量化自 Ornith-1.0-35B
Apple Silicon 上 Claude-Code 式本地编码代理
- 参数量
- 35B (3B 活跃)
- 商用
- 需自查
当前运行配置
nathansutton/Ornith-1.0-35B-UD-Q2_K_XL-MLX- 显存
- ~12GB 2-bit(估算)
查看详情查看运行方式
uvx --from git+https://github.com/nathansutton/mlxcc chad- 国内可达
- 需代理
当前运行配置
mlx-community/Qwen3.5-9B-OptiQ-4bit- 显存
- 暂无估算
查看详情查看运行方式
mlx_lm.generate --model mlx-community/Qwen3.5-9B-OptiQ-4bit- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- apache-2.0
- 上下文
- 128k
- 国内可达
- 需代理
当前运行配置
manjunathshiva/Qwen3.6-35B-A3B-tq3a-tqTe-g64- 显存
- 暂无估算
查看详情查看运行方式
pip install 'turboquant-mlx-full>=0.12.3' && python -m turboquant_mlx.generate --model manjunathshiva/Qwen3.6-35B-A3B-tq3a-tqTe-g64- 为何暂无估算
- 格式线索互相矛盾(含估算映射未支持的位宽),无法估算显存
- 许可证
- Apache-2.0
- 国内可达
- 需代理
Gemma4-Gutenberg-31B-Heretic-mlx-8Bit (ailexleon)
量化自 Gemma4-Gutenberg-31B-Heretic
MLX 8-bit 量化英文创意写作模型,适合故事小说生成
- 参数量
- 31B
- 商用
- 可商用
当前运行配置
ailexleon/Gemma4-Gutenberg-31B-Heretic-mlx-8Bit- 显存
- ~34GB 8-bit(估算)
查看详情查看运行方式
mlx_lm.generate --model ailexleon/Gemma4-Gutenberg-31B-Heretic-mlx-8Bit --prompt 'hello'- 许可证
- apache-2.0
- 国内可达
- 需代理
Qwen3.6-35B-A3B-OptiQ-4bit (mlx-community)
量化自 Qwen3.6-35B-A3B
Apple Silicon 4-bit MLX混精量化+MTP投机解码
- 参数量
- 35B-A3B
- 商用
- 可商用
aro-coder-4bit (ARO-Lang)
LoRA · 基于 Qwen3-Coder-30B-A3B-Instruct-4bit
为ARO语言微调的代码生成器,4bit量化,供ARO DSL开发者使用
- 参数量
- 30B (3B active)
- 商用
- 可商用
当前运行配置
prism-ml/Bonsai-8B-mlx-1bit- 显存
- ~1.2GB 1-bit(估算)
查看详情查看运行方式
pip install mlx-lm; pip install mlx@git+https://github.com/PrismML-Eng/mlx.git@prism; from mlx_lm import load; load('prism-ml/Bonsai-8B-mlx-1bit')- 许可证
- Apache-2.0
- 上下文
- 65k
- 国内可达
- 需代理