SWE-Universe: 百万级真实可验证 SWE 环境
Building agent(mini-sweagent + Qwen-Next-80A3)+ iterative self-verification + in-loop hacking 检测,从 33.3M GitHub PR(Pull Request,GitHub 上的代码合并请求)构造 807,693 实例,跑出 50 万轨迹 / 30B tokens,让 Qwen3-Max-Thinking 在 SWE-Bench Verified 达到 75.3%。标题中的 SWE = Software Engineering(软件工程)——指基于真实软件仓库的工程任务;SWE-Bench 是该方向的权威基准(见 SWE-bench 入门)。
概览
论文精读
3 难:低产率 / 弱 verifier / 高成本
SWE-Bench 已证明「真实仓库 PR = 自包含训练环境」,但手工扩展到百万级撞 3 类工程难题。SWE-Universe 论文给出一个全自动 pipeline 同时解决。
- 低产率:真实仓库依赖复杂,build 成功率 < 10%(很多 PR 装不上)
- 弱 verifier:浅薄 evaluation.sh 骗训练信号(用未来 patch 反向验证才能识破)
- 高成本:每实例 1 个大模型构造 N 次,80 万实例 = 几百万美元
- 规模目标:从 33.3M GitHub PR → 80 万+ 可验证 SWE 实例
PR 抓取 → Building agent → 5 scaffold rollout → rejection
完整 4 阶段 pipeline:33.3M GitHub PR → 启发式过滤到 1M 候选 → 自动构造 + 5 scaffold rollout → rejection sampling 收 50 万高质量轨迹。
- 阶段 1:33.3M PR → 启发式过滤 → 1M 候选
- 阶段 2:Building agent(mini-sweagent + Qwen-Next-80A3)自动构造 + test_patch / fix_patch 分离 + 迭代 self-verification
- 阶段 3:5 scaffold rollout(OpenHands / Mini / OpenHands / Claude-Code / Qwen-Code),跑出 50 万成功轨迹
- 阶段 4:rejection sampling 30B tokens + 拒绝浅薄 verifier
iterative self-verification + in-loop hacking detection
两个创新解决「弱 verifier」难题:1)自我验证(switch-to-bug / switch-to-resolved 双向)让构建 agent 反复测试;2)in-loop hacking 检测识别「跑未来 patch 才过、跑当前 patch 反而挂」的浅薄 verifier。
- iterative self-verification:构建 agent 生成 evaluation.sh 后跑 2 步(switch-to-bug / switch-to-resolved)
- in-loop hacking detection:检测「未来 patch 反而让 verifier 挂」的浅薄脚本(用真实测试 patch 反向验证)
- 5 scaffold rollout 提供跨环境/语言/格式的轨迹多样性
- Qwen3-Max-Thinking 最终 SWE-Bench Verified 75.3%(达开源模型 SOTA(State of the Art,当前最强水平))
807,693 实例 / 50 万轨迹 / 30B tokens / SWE-Bench 75.3%
从 33.3M GitHub PR 构造 80 万+ SWE 实例(75.9% non-hack success + 90,571 合成),跑出 50 万高质量轨迹,最终让 Qwen3-Max-Thinking 在 SWE-Bench Verified 达到 75.3%。
- 规模:807,693 实例(717,122 real + 90,571 synthesized 多语言)
- 质量:50 万成功轨迹 / 30B tokens(5 scaffold 各跑一轮)
- 终态:Qwen3-Max-Thinking 在 SWE-Bench Verified 75.3%(达开源 SOTA)
- 训练:256K 序列 mid-training(无 loss mask)+ agentic RL