Reddit r/LocalLLaMA (开源AI)· /u/Adorable-Cost-3249·· 2 小时前AI 评分49
实测:Qwen3.8-27B Q4_K_M 在单张 RTX 3090 上搭配 OpenCode 的吞吐与编码任务表现
Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure
AI 导读
作者在单张 RTX 3090 24GB 上用 llama.cpp 和 OpenCode 2.0.20 本地运行 Qwen3.8-27B Q4_K_M,131,072 token 上下文下生成速度 20.9–36.4 tok/s,缓存续算首 token 低至 0.46 秒。
来源:Reddit r/LocalLLaMA (开源AI) · reddit.com