跳到正文
原文
Reddit r/MachineLearning· /u/heyitsdannyle·· 23 小时前AI 评分39

SWE-Race:基于 188 个真实并发 bug 的编程智能体基准,三款模型成绩出炉

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

AI 导读

SWE-Race 基准从约 100 个 Python 项目的已合并 PR 中提取真实并发 bug 构建任务,由项目自身测试在无网络容器中评分。单次尝试下 GLM-5.3 Flash 得分 85%,两到三次尝试后为 82%,与 GPT-5.6 Luna(81%)在误差范围内接近;硬任务上三个模型分别为 50%、45% 和 23%。

来源:Reddit r/MachineLearning · reddit.com