SWE-agent 执行引擎拆解:一次任务如何走完 N 轮对话
前面三页分别给了 Agent 循环的心智模型、六部件规格和动手实现;本页换一个问法:这些抽象机制在一线开源实现里长什么样? 标本是 Princeton SWE-agent——NeurIPS 2024 论文(arXiv:2405.15793)+ 本地开源代码(v1.1.0,commit 3ea751c)。我们顺着一次任务的生命周期逐层拆开:提交任务 → 每轮模型返回一个工具调用 → 多步执行 → 直到 submit 终止,看清 Agent 在抽象模型层的工作原理。
概览
| 项目 | 说明 |
|---|---|
| 本页定位 | 源码级案例拆解:把"轮次 / 工具调用 / 上下文 / 终止"四个抽象概念逐个指认到 SWE-agent 的真实代码行,回答"Agent 执行引擎最少需要哪些机制" |
| 标本版本 | 论文:arXiv:2405.15793v3(NeurIPS 2024 camera-ready);代码:sweagent 1.1.0(commit 3ea751c,2026-07-16,1.x 架构)——论文讲设计动机,代码讲当前实现,两者不一致处以代码为准并注明 |
| 前置知识 | 读过 agent-harness-loop.html(Agent = harness + LLM)或 agent-minimal-six-parts.html(六部件规格);知道"API 发消息进、收消息出"即可 |
| 读完会什么 | 能讲清一轮(step)的四拍结构;能说出模型每轮实际看到什么(history + 模板);能解释 submit 终止的哨兵机制与全部 12 种退出状态;理解 ACI(Agent-Computer Interface)为什么比模型本身更决定成绩 |
| 不讲什么 | 协议字段细节(见 协议讲义 与第 3 层字段手册)、SWE-bench 评分链路(见 swe-bench-evaluation.html)、轨迹分析方法(见 1.6 过程分析组) |
§ 1 · 全景:一次任务的一生(setup → N × step → 提交)
一次任务的完整时序
用户执行 sweagent run 后,RunSingle(sweagent/run/run_single.py)只做六件事:启动环境 → 调用 agent.run() → 保存预测 → 关环境。真正的"Agent"全部在 agent.run() 里。下图把一次任务从头到尾画出来——注意循环里每一轮都恰好包含一次模型调用和一个环境动作:
- setup(一次性):把工具 bundle 安装进容器 /root/tools/;把问题文本注入容器环境变量 PROBLEM_STATEMENT;按序向 history 写入 system 模板 → 示范(可选)→ instance 模板三条种子消息(agents.py:561-606,写历史在 603-605 行)
- loop(N 轮):每轮 = 一次模型查询 + 一个动作 + 一条观测,每轮结束立刻把轨迹落盘——崩溃也能复盘到最后一轮
- after(一次性):save_trajectory 写 <instance_id>.traj(顶层 5 键:trajectory / history / info / replay_config / environment);save_predictions 从中抽出 submission 写成 SWE-bench 兼容的 .pred(model_name_or_path / instance_id / model_patch)交给评测
主循环全貌:一个 while,五行代码
剥掉所有细节后,SWE-agent 的"执行引擎"就是下面这段代码(DefaultAgent.run,逐字摘录)——整个循环的唯一退出条件是 step_output.done 为真:
self.setup(env=env, problem_statement=problem_statement, output_dir=output_dir)
# Run action/observation loop
self._chook.on_run_start()
step_output = StepOutput()
while not step_output.done:
step_output = self.step()
self.save_trajectory()
self._chook.on_run_done(trajectory=self.trajectory, info=self.info)
教学上值得停一秒:没有 for 循环、没有 max_steps、没有"规划器"——一个条件循环加上每步落盘,就是工业级 Agent 引擎的骨干。复杂性全部藏在 step() 里,那是 § 2-§ 6 要拆的内容。
§ 2 · 轮次(step):引擎的最小节拍
step():一轮的完整定义
assert self._env is not None
self._chook.on_step_start()
n_step = len(self.trajectory) + 1
self.logger.info("=" * 25 + f" STEP {n_step} " + "=" * 25)
step_output = self.forward_with_handling(self.messages)
self.add_step_to_history(step_output)
self.info["submission"] = step_output.submission
self.info["exit_status"] = step_output.exit_status
self.info.update(self._get_edited_files_with_context(patch=step_output.submission or ""))
self.info["model_stats"] = self.model.stats.model_dump()
self.add_step_to_trajectory(step_output)
self._chook.on_step_done(step=step_output, info=self.info)
return step_output
读法:forward_with_handling(self.messages) 完成"查询模型 + 解析 + 执行"(含格式错误重试,§ 3);随后三件事记账——观测写回 history、info 更新(提交 / 退出状态 / 编辑文件 / 成本统计)、这一步追加进 trajectory。轮次号 n_step 直接取轨迹长度 +1。
| 概念 | 定义 | SWE-agent 里的对应物 |
|---|---|---|
| 轮次 / step | 一次"查询 → 解析 → 执行 → 观测"的完整节拍 | trajectory 数组的长度;每步向 history 追加 2 条消息(assistant 原始输出 + user 观测) |
| 消息 / message | API 视角的一条 role/content | history 列表的元素:3 条种子 + 每轮 2 条(不含重试时的临时错误消息) |
| 动作 / action | 一轮中恰好一个命令 | function calling 模式下强制"恰好一个 tool call",多了少了都报格式错误(parsing.py:441-454) |
| 轨迹步 / TrajectoryStep | 落盘的单轮记录 | 8 字段:action / observation / response / state / thought / execution_time / query / extra_info(types.py:44-52);其中 query 是本步发给模型的完整消息快照 |
解决的实例中位数 12 步 / $1.21,未解决的均值 21 步 / $2.52;93% 的解决实例在耗尽预算前主动 submit。作者据此总结:
这也解释了为什么 v1.1.0 代码里根本没有"最大步数"这个配置(全仓 grep 实证):步数不是好的预算单位——一轮可能花 500 token 也可能花 5 万 token。真正的护栏是钱、时间、连续错误三类,见 § 6。
§ 3 · 一步的内部:查询 → 解析 → 执行 → 记账
forward():五行走完"问模型 → 拆输出 → 执行"
output = self.model.query(history)
step.output = output["message"]
step.thought, step.action = self.tools.parse_actions(output)
step.thinking_blocks = output.get("thinking_blocks", [])
if output.get("tool_calls") is not None:
step.tool_call_ids = [call["id"] for call in output["tool_calls"]]
step.tool_calls = output["tool_calls"]
self.logger.info(f"💭 THOUGHT\n{step.thought}\n\n🎬 ACTION\n{step.action.strip()}")
self._chook.on_actions_generated(step=step)
return self.handle_action(step)
四拍各自的位置:①查询 model.query(history)——发给模型的是"完整历史经处理器链变换后"的消息(§ 4);②解析 parse_actions——把模型输出拆成(思考,动作)二元组;③执行 handle_action(step)——动作进容器跑,观测写回 step(§ 5);④记账在 step() 里完成(§ 2)。一轮恰好一次模型调用——模型不知道执行结果之前不会再被问第二次。
本步唯一一次模型调用"] Q --> P{"parse_actions
解析成功?"} P -- "否:FormatError / 黑名单 / bash 语法错误" --> R{"重试次数 < 3?"} R -- 是 --> ER["错误提示模板 + 出错输出
临时拼进 history"] --> Q R -- 否 --> X1["exit_format 退出
(先自动抢救提交,§6)"] P -- 是 --> H["handle_action(action)"] H --> B{"动作是 exit / 命中黑名单?"} B -- 是 --> X2["exit_command 退出 / 拒绝并重试"] B -- 否 --> C["env.communicate()
容器持久 bash 执行"] C --> T{"观测含特殊标记?"} T -- "RETRY token(lint 失败等)" --> Q T -- "FORFEIT token" --> X3["exit_forfeit:自动提交后退出"] T -- "SUBMISSION 哨兵" --> D["读 /root/model.patch
done = True"] T -- 无 --> O["观测经模板写回 history
(超长截断 10 万字符)"] O --> SAVE["轨迹落盘 save_trajectory()"] D --> SAVE
模型的"工具调用"怎么落地:两条解析路线
模型不会"调用"任何东西——它只输出文本。把文本变成一个可执行动作是解析器的活。SWE-agent 两代默认配置恰好各代表一条路线,这也是本站反复出现的那条"协议轴":
| 维度 | ThoughtActionParser(论文时代 / 0.7 配置) | FunctionCallingParser(1.x 默认) |
|---|---|---|
| 动作载体 | 输出中最后一个顶层三反引号代码块就是动作 | API 原生 tool_calls 字段 |
| 思考部分 | 代码块之外的全部文本都算 thought | message 文本本身 |
| 硬约束 | 找不到代码块 → 抛 FormatError 重试 | 恰好一个 tool call——0 个或多个都报错(missing / multiple) |
| 工具文档 | 内嵌在 system 模板的 {{command_docs}} 里 | 工具 schema 走 API 的 tools 参数 |
| 配置出处 | config/sweagent_0_7/07.yaml(parse_function: thought_action) | config/default.yaml(type: function_calling) |
为什么论文时代故意选纯文本?模型无关(任何能输出 markdown 的模型都能跑)、协议即提示词(改动不用动 API 参数)、可截断可重放(历史里每条消息都是纯文本)。这正是 ollama-tool-call-substitute.html 里"协议上移到提示词"路线的学术版先例。
code_block_pat = re.compile(r"^```(\S*)\s*\n|^```\s*$", re.MULTILINE)
stack = []
last_valid_block = None
for match in code_block_pat.finditer(model_response["message"]):
if stack and not match.group(1): # Closing of a code block
start = stack.pop()
# Check if it's not nested within another block
if not stack:
last_valid_block = (start, match)
elif match.group(1) is not None: # Opening of a code block
stack.append(match)
if last_valid_block:
start, end = last_valid_block
thought = model_response["message"][: start.start()] + model_response["message"][end.end() :]
return thought, model_response["message"][start.end() : end.start()]
msg = "No action found in model response."
raise FormatError(msg)
设计巧思:取最后一个顶层代码块而不是第一个——模型可以在思考里放心引用代码片段(包在 ``` 里),只要最后一个代码块是给环境的命令即可。提示词工程("RESPONSE FORMAT 约定")与解析器互为表里:提示词负责让模型大概率守约定,解析器负责在小概率违约时抛错重试。
§ 4 · 模型看到什么:history 与模板系统
消息序列的骨架:三条种子 + 每轮两条
setup 阶段写入三条种子消息,之后每轮追加"assistant 原始输出 + user 观测"两条,形成严格交替:
#0 system system_template ← 一句话角色设定
#1 user demonstration(可选)← 一条完整示范轨迹
#2 user instance_template ← 任务书:仓库位置 + PR 描述 + 五步指示
#3 assistant 模型第 1 轮原始输出(thought + action)
#4 user next_step_template ← "OBSERVATION:\n{{observation}}"
#5 assistant 模型第 2 轮原始输出
#6 user next_step_template ← 观测
…(交替直到 submit)
system_template: |-
You are a helpful assistant that can interact with a computer to solve tasks.
instance_template: |-
<uploaded_files>
{{working_dir}}
</uploaded_files>
I've uploaded a python code repository in the directory {{working_dir}}.
Consider the following PR description:
<pr_description>
{{problem_statement}}
</pr_description>
…(下接五步指示:找代码 → 写复现脚本 → 改源码 → 重跑复现 → 想边界情况)
next_step_template: |-
OBSERVATION:
{{observation}}
注意两个细节:① 1.x 的 system 模板只剩一句话——工具文档改由 API 的 tools 参数承载(论文时代则内嵌 {{command_docs}} 全文);② instance 模板里"I've already taken care of all changes to any of the test files… you DON'T have to modify the testing logic"是提示词级约定,不是代码硬护栏——引擎并不阻止模型改测试文件,这一点常被误传。
观测怎么进历史:三种模板 + 10 万字符截断
执行结果不是原文塞回——按内容分三路处理(add_step_to_history):
- 空输出 → 替换成固定话术 "Your command ran successfully and did not produce any output."——防止模型把"没输出"误读为"没执行"
- 超长输出(> max_observation_length,默认 100,000 字符)→ 截断 + 标注省略了多少字符
- 正常输出 → 包进 next_step_template(OBSERVATION: 前缀)
真正发给模型前,history 还要现算一道"处理器链"(agents.py:540-551)。论文时代的经典配置是 last_n_observations(n=5):只保留最近 5 条完整观测,更早的折叠成一行 Old environment output: (n lines omitted);1.x 默认换为 cache_control(last_n_messages=2)(Anthropic prompt caching 标记)。处理器共 7 种可选(history_processors.py:390-399),恒等的 default 也在其中。
每次发送前本地数 token(models.py:690-703),超过模型最大输入窗口直接抛 ContextWindowExceededError 退出——不做摘要、不做折叠,把抢救交给 § 6 的自动提交。这是刻意的简单:摘要意味着又一次模型调用与新的失败面,学术 baseline 选择把复杂性让渡给可复现性。
§ 5 · 动作在哪里执行:环境、持久 bash 与观测的形成
communicate():字符串进,字符串出
rex_check = "silent" if check else "ignore"
r = asyncio.run(
self.deployment.runtime.run_in_session(BashAction(command=input, timeout=timeout, check=rex_check))
)
output = r.output
…
return output
动作字符串经 run_in_session(BashAction(...)) 送进容器里同一个持久 bash 会话——所以 cd、环境变量、后台进程在轮次之间是连续的(模型第 3 轮 cd 进的目录,第 8 轮还在)。返回值是 stdout + stderr 合并的原始文本,环境层不加任何包装。
- "Your command ran successfully…"这类话术不是环境加的,是 agent 层的 next_step_no_output_template(§ 4)——环境只回原始字节,"怎么向模型转述"是引擎的事
- 同理,截断也发生在 agent 层(10 万字符),环境层只按 execution_timeout(默认 30 秒)掐表
- 工具 bundle 可声明 state_command:每步执行后在容器内写一个 /root/state.json(如当前工作目录),agent 读出后并入模板变量——instance_template 里的 {{working_dir}} 就是这么来的,Python 代码里并不存在这个变量,纯靠这条 side-channel
- 单命令超时 30 s;连续 3 次超时直接判死退出(exit_command_timeout,tools.py:139-151 / agents.py:968-975)
- 单实例累计执行时长上限 1800 s(exit_total_execution_time)
- 多行命令先经 guard_multiline_input 改写为 heredoc 再下发(agents.py:961),避免交互式 bash 把多行输入拆乱
§ 6 · 终止设计:submit 不是 API,是一个 bash 脚本
submit 工具的全部实现:15 行 bash
main() {
cd $ROOT
# Check if the patch file exists and is non-empty
if [ -s "/root/test.patch" ]; then
# Apply the patch in reverse
git apply -R < "/root/test.patch"
fi
git add -A
git diff --cached > /root/model.patch
echo "<<SWE_AGENT_SUBMISSION>>"
}
main "$@"
值得逐行讲:① 先把评测用的 gold 测试补丁反向移除(git apply -R),保证导出的补丁只含模型自己的改动;② git add -A && git diff --cached 把工作区全部改动导出为 /root/model.patch——提交物就是一份标准 diff;③ 打印哨兵字符串 <<SWE_AGENT_SUBMISSION>>。引擎侧在每条观测里扫描该哨兵(tools.py:372-376),命中即读回补丁、置 done=True(agents.py:883-904)。终止信号与观测数据共用同一文本通道——不需要任何带外协议。
对照姊妹页的抽象规则:"模型回复纯文本 = 对话结束"。SWE-agent 把"结束"也变成一个显式工具调用:模型决定任务完成的时机,引擎负责把决定翻译成可评测的产物(补丁)。
轮次何时结束整个任务:全部 exit_status 一览
循环出口只有两类——主动终局(模型自己说"做完了 / 不做了")与护栏熔断(预算或错误触线)。代码里穷尽的退出状态共 12 种:
| 类别 | exit_status | 触发机制 | 代码位置 |
|---|---|---|---|
| 主动终局 | submitted | 观测含 <<SWE_AGENT_SUBMISSION>> 哨兵,读 /root/model.patch 作为提交 | agents.py:883-904 |
| exit_command | 模型输出的动作恰为 exit | agents.py:949-956 | |
| exit_forfeit | 模型调用 forfeit 工具(打印 ###SWE-AGENT-EXIT-FORFEIT###)——认输也先自动提交再退 | agents.py:1001-1002 | |
| 预算护栏 | exit_cost | 实例花费 > per_instance_cost_limit(默认 $3)或调用数超限 | models.py:660-670 |
| (全局穿透) | 全部实例总花费 > total_cost_limit → 异常穿透,整个 batch 停止 | models.py:655-658 | |
| exit_context | 输入 token 数超过模型窗口(发送前本地数 token) | models.py:690-703 | |
| exit_total_execution_time | 单实例累计执行 > 1800 s | agents.py:1018-1019 | |
| exit_command_timeout | 连续 3 次命令超时 | agents.py:968-975 | |
| 错误退出 | exit_format | 格式 / 黑名单 / bash 语法错误连续重试 3 次仍失败 | agents.py:1106-1107 |
| exit_api | API 调用 tenacity 重试 20 次耗尽 | agents.py:1187-1192 | |
| exit_environment_error / exit_error | 环境异常 / 未知异常 | agents.py:1193-1210 |
所有"错误退出"都先过 attempt_autosubmission_after_error(agents.py:823-868):只要容器还活着,就手动执行 git add -A && git diff --cached > /root/model.patch 尽力抢救一份补丁再退出。没有"白跑"的实例——做到哪算哪,现状即提交。论文的 $4 预算实验同样如此:超预算时自动提交已有改动。
§ 7 · 为什么接口比模型更决定成绩:ACI 与消融证据
ACI:Agent 与计算机之间的抽象层
类比 HCI(人机交互):人类工程师受益于 IDE,而语言模型是一类全新的"终端用户",有自己的能力与需求,理应拥有专门为其打造的界面。ACI 规定两件事:模型可用哪些命令、环境状态如何回传给模型。论文给出四条设计原则:
- ① 动作要简单好懂——少量选项 + 简洁文档,别让模型在工具列表里迷路
- ② 动作要紧凑高效——重要操作合并进尽可能少的动作,一步取得实质进展(反例:逐条翻搜索结果的"迭代式"接口)
- ③ 环境反馈要有信息量但简洁——告知状态与动作效果,但不冗余(文件查看器一次只给 100 行窗口)
- ④ 护栏缓解错误传播、加速恢复——语法检查器自动拦下坏编辑,让模型立刻重试
每个设计决定都有价格标签
论文 Table 3 的消融实验,把 ACI 的每个部件单独拆掉再测——成绩掉多少,就是该部件的价值:
| 配置变体 | %Resolved | Δ | 结论 |
|---|---|---|---|
| 完整配置(100 行窗口 + 摘要式搜索 + 最近 5 条观测 + lint 护栏) | 18.0 | — | 基准 |
| 去掉编辑后的 lint 护栏 | 15.0 | −3.0 | 坏编辑被静默接受,错误一路传播 |
| 保留全部历史(不折叠旧观测) | 15.0 | −3.0 | 上下文冗余直接伤成绩 |
| 文件查看器窗口缩到 30 行 | 14.3 | −3.7 | 看得太少,上下文不足 |
| 文件查看器一次给整个文件 | 12.7 | −5.3 | 看得太多,淹没关系信息 |
| 搜索改成逐条翻页(类 Vim / VSCode) | 12.0 | −6.0 | 比无搜索(15.7)还差——agent 穷尽翻页耗尽预算 |
| 无编辑接口(只能整文件重写 / sed) | 10.3 | −7.7 | 低效且易错,且缺乏编辑后反馈 |
GPT-4 Turbo 的 2294 条轨迹中,51.7% 至少出现过 1 次失败编辑;单次编辑成功率 90.5%,但失败过一次后降到 57.2%——错误会自我繁殖。lint 护栏(编辑后立刻语法检查,不合格即丢弃并要求重试)斩断的正是这条链:"Invalid edits are discarded, and the agent is asked to try editing the file again."(论文 §3)
§ 8 · 收口:最小伪代码与组件映射
执行引擎的最小重建(每行对应真实代码位置)
env.start(); tools.install(env) # ① 起容器、把工具脚本装进 /root/tools
history = [system, demo?, instance] # ② 三条种子消息(Jinja2 模板渲染)
while not step.done: # ③ 唯一循环条件(agents.py:1284,没有步数上限)
messages = 处理器链(history) # ④ 折叠旧观测 / 打缓存标记(§4)
for retry in range(3): # ⑤ 格式错误最多重试 3 次
output = model.query(messages) # ⑥ 本步唯一一次模型调用;超钱/超窗抛异常退出
thought, action = parse(output) # ⑦ 恰好解析出一个动作(``` 块 或 一个 tool call)
if 解析失败: messages += 错误提示; continue # ⑧ 错误消息只用于重试,不进正式历史
break
observation = env.communicate(action) # ⑨ 送进容器持久 bash,拿回 stdout+stderr
if "<<SWE_AGENT_SUBMISSION>>" in observation: # ⑩ submit 脚本打印的哨兵
step.submission = 读 /root/model.patch; step.done = True
history += [assistant(output), user(观测模板)] # ⑪ 写回历史(空输出换话术 / 10 万字符截断)
trajectory.append(step); save_trajectory() # ⑫ 每步落盘 .traj,崩溃可复盘
save_predictions(submission) # ⑬ 抽出补丁写 .pred,交给 SWE-bench 评测
拿着这段伪代码去仓库里按图索骥,每一行都能落到 agents.py / models.py / tools/ 的具体位置——这就是"抽象模型层原理"与"工程实现"之间最短的距离。
抽象部件 ↔ SWE-agent 真身对照表
| 六部件规格(agent-minimal-six-parts) | SWE-agent 1.1.0 真身 | 本页位置 |
|---|---|---|
| ① 工具清单 | tools/ bundle 目录(config.yaml 声明 + bin/ 脚本);function calling 下每个命令转 OpenAI schema(commands.py:133-160) | § 3 |
| ② 系统提示词 | config/default.yaml 模板组(system / instance / next_step) | § 4 |
| ③ 对话历史 | self.history + 处理器链(查询前现算 messages) | § 4 |
| ④ 模型调用器 | LiteLLMModel.query:20 次重试 + litellm 成本记账 + 发送前数 token 超窗守卫 | § 3 / § 6 |
| ⑤ 响应解析与调度 | tools.parse_actions(可插拔解析器)+ handle_action | § 3 |
| ⑥ 主循环 | DefaultAgent.run 的 while not done(每步落盘) | § 1 / § 2 |
| 终止 / 护栏(工程视角多出) | submit 哨兵工具 + 12 种 exit_status + autosubmission | § 6 |
- 1.0 重设计(2025-02):执行层外包给独立包 SWE-ReX、工具改为可组合 bundle、模型层统一走 litellm——"codebase has been nearly rewritten from scratch",Agent 类大幅简化(官方 changelog)
- mini-swe-agent:官方文档首页已挂出"SWE-agent has been superseded by mini-swe-agent… maintenance-only mode"横幅——后者是约 100 行、只用 bash 一个工具、完全线性历史的极简实现,官网宣称 SWE-bench Verified > 74%(2025-11 新闻口径,含 Gemini 3 Pro)。本专题 topics/01 的 90 行 mini_agent 与它是同一条极简路线的教学版
- 如果没有哨兵字符串,改用"模型返回里不再含工具调用"作为终止信号,会丢掉什么?(提示:forfeit 与 submit 的区分、补丁导出的时机、轨迹可审计性)
- v1.1.0 用"钱 + 时间 + 连续错误"替代"最大步数"做预算——在你的场景里,哪个物理量最贴近真实约束?
- § 7 的消融里"逐条翻页搜索"比"没有搜索"更差——这对你设计工具返回格式有什么启示?