跳到正文
原文
Reddit r/LocalLLaMA (开源AI)· /u/Prestigious-Taste-63·· 5 小时前AI 评分38

从零训练 3.87B 参数 MoE 模型 Apex-2,仅用 86.5B tokens 预训练

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

AI 导读

Reddit 用户从零训练了纯 MoE 模型 Apex-2,总参数 3.87B、每 token 激活 1.45B,预训练仅用 86.5B tokens。SFT 后 HumanEval+ 得 41.5,与用了 18T tokens 的 Qwen2.5-1.5B 相当,但 MMLU 28.6、数学和知识仍差距明显。

来源:Reddit r/LocalLLaMA (开源AI) · reddit.com