跳到正文
原文
Reddit r/LocalLLaMA (开源AI)· /u/Designer_Elephant227·· 2 小时前AI 评分20

单张 r9700 跑 Qwen3.8-Flash-Next exl3:230k 上下文下 prefill 863 t/s、decode 35 t/s,速度正常吗?

Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

AI 导读

一位用户在单张 r9700 上用 exllamav3 rocm 分支运行 Qwen3.8-Flash-Next(exl3 5.05 bpw),262144 上下文、q8 KV cache,112 个专家在 GPU、400 个在 CPU,开启 mtp drafting(接受率约 53-54%),102gb bf16 ngram 表从 nvme 流式读取。

来源:Reddit r/LocalLLaMA (开源AI) · reddit.com