Reddit r/LocalLLaMA (开源AI)· /u/Specific-Tax-6700·· 3 小时前AI 评分52
AgrillaMoE:在 16GB GPU 上以 2-bit 量化运行 Qwen3.6-35B-A3B 并支持 MoE 扩展
poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding
AI 导读
作者将 llama.cpp 服务端 fork 为 AgrillaMoE,专用于 Qwen3.6-35B-A3B 配合 Unsloth 量化。在租用的 V100 16GB 上用 UD-Q2_K_XL 2-bit 量化达到约 57-60 tok/s,兼容 OpenAI 和 Anthropic API,可直接对接 Claude Code。
来源:Reddit r/LocalLLaMA (开源AI) · reddit.com