跳到正文
原文
Reddit r/LocalLLaMA (开源AI)· /u/theexile1337·· 3 小时前AI 评分40

实测对比:RTX 5090 上 ninfer 跑 Qwen3.8 27B 比 llama.cpp 快 2.3 倍,4 种配置同提示词测速与质量

2.3x faster Qwen3.8 27B on a 5090: ninfer vs llama.cpp, 4 setups, same prompt - speed and quality tested

AI 导读

网友在 RTX 5090 上用 Qwen3.8 27B(thinking on xhigh、120k context)对比 ninfer 与 llama.cpp 的 4 种配置:ninfer 非 NVFP4 版解码约 147 tokens/s,而 llama.cpp Q4_K_M 无 MTP 仅约 68 tokens/s,开启 MTP 后可达约 141 tokens/s。

来源:Reddit r/LocalLLaMA (开源AI) · reddit.com