门户首页
Agent 开发 · 心智模型

理解 Agent 开发过程:harness + LLM 的任务循环

很多人以为 Agent 是某种"更聪明的 AI",能自动完成复杂任务。真相是:所谓 Agent,就是一段循环代码(harness)+ 一个外部大语言模型(LLM,Large Language Model)协同工作的产物。 这一页把"AI 自动做事"的黑盒拆开——从 30 秒 LLM API 入门,到 harness 6 大组件、主循环伪代码、SWE-bench S1-S7 真实例子、再到一个 30 行可运行 harness——让你下次再听到"Agent"时不带神秘感。

生成时间:2026-09-02 · 版本 v0.2 · 生成 Agent:MiniMax Code (LLM: MiniMax-M3) · 载体:agentsoft-research-platform teaching-web-platform

概览

1
核心等式
6
harness 组件
1
主循环
4
主流模式
项目说明
本卡定位Agent 开发入门页 · 心智模型层 · 7 张卡片从零讲到实战
前置知识会 Python 基础语法即可;不要求调过 LLM API(卡片 1 铺垫)
读完会什么能讲清"Agent 是什么";能照卡片 7 写一个 30 行最小 harness;能看懂 SWE-bench S1-S7 的 orchestration 是怎么对应到 harness 的
读完去哪儿llm-protocol-fundamentals.html(协议层)/ llm-api-schema-reference.html(Schema 层)

第一讲 · 30 秒 LLM API 入门(零基础铺垫)

第一讲 · LLM API 是什么(1 概念 + 1 cURL + 1 Python)
如果你是第一次接触 LLM API,这一节让你 30 秒内建立"发 JSON 进、收 JSON 出"的心智模型;如果已经熟悉,可以跳到卡片 2。
第 1 讲 · 1 概念 + 0 实验

LLM API:一次调用 = 发 JSON 进、收 JSON 出

大模型本身是一个 HTTP 服务。你按它规定的 JSON(JavaScript Object Notation,一种人类可读的文本数据格式)格式发请求,它按规定的 JSON 格式回响应。这个 JSON 格式就是协议(llm-protocol-fundamentals.html 专门讲)。理解 LLM API 的关键:它就是"一次请求一次响应",不"记住"任何前文,没有"自动做事"的能力——你看到的"Agent 自动完成多步任务",全是外面包了一层循环代码(那就是 harness,下一节讲)。这里 LLM API 指"通过 HTTP 接口调用大模型"——API(Application Programming Interface,应用程序接口)即程序之间约定的调用方式。

最小 cURL 调用(5 秒看明白)
curl https://api.openai.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "1+1=?"}]
  }'

# → 响应(简化)
{
  "choices": [{
    "message": {"role": "assistant", "content": "1+1 等于 2。"},
    "finish_reason": "stop"
  }]
}
等价 Python(3 行 SDK 调用)
from openai import OpenAI
client = OpenAI()                                    # 从环境变量读 API key
r = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "1+1=?"}],
)
print(r.choices[0].message.content)                  # → "1+1 等于 2。"
关键认知:一次 LLM API 调用 = 单回合。你发 messages 数组(包含完整历史),它回 content + finish_reason。它不"记住"上一轮——多轮对话全靠你把历史拼回去。所以"Agent 能连续做事"这件事不是 LLM 的能力,是外面包的循环代码的功劳。
从"单调用"到"自动做事":缺什么?
  • ❌ 没有持久状态:每次调用都得把全部历史塞进 messages,否则 LLM 不知道之前说了啥
  • ❌ 不能调外部工具:默认只会"说",不能"做"(查数据库、改文件、跑命令)
  • ❌ 不知道何时停:复杂任务需要"反复调 LLM → 看结果 → 决定下一步",LLM 自己不知道什么时候该收手
  • ❌ 不处理错误:API 超时、JSON 解析失败、工具执行报错——全得有人兜底

把这 4 件事补齐的"那一层"——就是 harness。下一节破题。

第二讲 · 破题:Agent = harness + LLM

第二讲 · 一个等式拆解黑盒(1 等式 + 1 图 + 1 心法)
把"AI 自动完成任务"的神秘感拆掉,看到每一个 ownership 边界。
第 2 讲 · 1 等式 + 1 图 + 0 实验

一个等式:Agent = harness(你写的代码)+ LLM(外部 API)

所有"Agent"——无论是 LangChain、AutoGPT、还是本仓 SWE-bench(Software Engineering benchmark,用真实 GitHub issue 评测代码 agent 的软件工程基准测试)的 S1-S7——拆到底都是同一个结构:一段你自己写的代码(harness),去调一个外部 LLM 服务(OpenAI / Anthropic / Ollama),中间穿插工具调用和状态管理。"Agent"不是一个新物种,是一种编程范式。下面这张组件图把 ownership 划清楚。

图 1 · Agent 的三块组件(ownership 边界一目了然)
flowchart LR subgraph Harness["🟦 HARNESS(你写的代码 · 完全可控)"] direction TB State["State Manager
messages[] / items[] 维护"] Tools["Tool Registry
工具定义 + 执行"] Parser["Output Parser
tool_calls / finish_reason 判别"] Loop["Loop Controller
while not done"] Err["Error Handler
重试 / 降级 / 中止"] Log["Logger
trajectory 记录"] end subgraph LLM["🟧 LLM(外部 API · 不可控)"] LLM1["按 messages + tools
生成下一轮回复"] end subgraph World["🟩 EXTERNAL WORLD"] W1["文件系统 / 数据库
Git / Shell / API"] end Loop -->|"构造 messages"| LLM1 LLM1 -->|"返回 content / tool_calls"| Parser Parser -->|"finish_reason = tool_calls"| Tools Tools -->|"执行副作用"| W1 W1 -->|"结果"| State State -->|"拼回 messages"| Loop Err -.->|"异常兜底"| Loop Log -.->|"全程旁路"| Loop
图 1 · Agent = harness(蓝)+ LLM(橙)+ 外部世界(绿)。你只对蓝色部分有完全控制权。
心智模型一句话:LLM 是函数,harness 是调用方。LLM 接收 messages + tools,返回 content + tool_calls;harness 决定何时调、怎么用结果、什么时候停、错了怎么办。Agent 的"智能"来自 LLM,Agent 的"行为"来自 harness。
一个反直觉的事实
  • LLM 不知道任务进度:它只看到当前这一轮的 messages,不知道"我正在做第 3 步 / 总共 10 步"
  • LLM 不知道工具的结果格式:你把工具输出塞进 role:"tool" 消息,它才"看见"
  • LLM 不知道历史轨迹的元信息:时间戳、token 用量、工具调用次数——这些都在 harness 的 logger 里
  • 所以 "AI 自动完成多步任务"这个体验 = harness 的循环 + LLM 的每步推理,缺一不可
spec §1 心智模型 paper: ReAct (ICLR 2023 · ICLR = International Conference on Learning Representations,机器学习顶会)

第三讲 · harness 6 大组件

第三讲 · 拆开 harness(6 组件 × 1 句 + 1 段最小代码)
把蓝色那块的内部结构摊开:每个组件做什么、用什么最小代码表达。
第 3 讲 · 6 组件 + 1 速查表

拆开 harness:6 大组件 + 1 张速查表

所有成熟的 Agent 框架(LangGraph、AutoGen、Claude Agent SDK、本仓 agent_runtime)拆到底都是这 6 个组件的排列组合。理解这 6 个,你就看穿了所有 Agent 框架的"魔法"。

① State Manager(状态管理器)

职责:维护 messages[] / items[],每次循环后追加新消息(assistant 回复、tool 结果)。这是 LLM 唯一能"看到历史"的窗口。

messages = [{"role": "user", "content": "帮我看看 README"}]
# 循环里:messages.append(assistant_msg); messages.append(tool_result_msg)
② Tool Registry(工具注册表)

职责:声明 LLM 能调的工具(tools=[...]),并提供实际执行函数(name → callable 的映射)。声明与实现解耦。

TOOL_REGISTRY = {
    "read_file": lambda p: open(p).read(),
    "bash":      lambda c: subprocess.run(c, shell=True, capture_output=True).stdout,
}
tools = [{"type": "function", "function": {"name": n, "description": "...", "parameters": {...}}} for n in TOOL_REGISTRY]
③ Output Parser(输出解析器)

职责:从 LLM 响应里抽出 content / tool_calls / finish_reason。手写 Agent 最常踩的坑都在这里:arguments 是 JSON 字符串需要 json.loads;流式 chunk 要按 index 聚合;finish_reason="tool_calls" 时 content 可能是 null。

msg = response.choices[0].message
if msg.tool_calls:
    for tc in msg.tool_calls:
        args = json.loads(tc.function.arguments)   # ← 字符串,不是对象!
        result = TOOL_REGISTRY[tc.function.name](**args)
elif msg.finish_reason == "stop":
    final_answer = msg.content
④ Loop Controller(循环控制器)

职责:while not done,附带三层守门:max_iterations(防死循环)/ wall_time(防超时)/ max_tokens(防上下文爆)。

for turn in range(MAX_ITERATIONS):                  # 守门 1:轮次
    if time.time() - start > MAX_WALL_TIME: break   # 守门 2:墙钟
    r = client.chat.completions.create(model=..., messages=messages, max_tokens=...)  # 守门 3:单次上限
    ...
⑤ Error Handler(错误处理器)

职责:网络超时 → 重试;tool_call_id 配错 → 校验抛错;上下文超限 → 摘要压缩;LLM 返回非法 JSON → 兜底默认值。没有错误处理,Agent 跑 10 次崩 8 次。

try:
    r = client.chat.completions.create(..., timeout=60)
except openai.APITimeoutError:
    time.sleep(2 ** retry_count)                    # 指数退避
    continue
except json.JSONDecodeError:                        # arguments 不是合法 JSON
    return fallback_default
⑥ Logger(轨迹记录器)

职责:全程旁路记录每轮的 messages、tool_calls、tool 结果、token 用量、墙钟。没有 logger,Agent 跑完你不知道它怎么做的——所有"为什么这次没成功"的事后分析都靠这个。本仓的 experiments/<id>/s4_worker.log + result.json 就是这一层。

log.write({
    "turn": turn, "ts": now(),
    "messages_snapshot": messages[-3:],            # 最近 3 条,避免日志爆炸
    "tool_call": msg.tool_calls, "tool_result": result,
    "usage": r.usage.total_tokens, "wall_time": time.time() - start,
})
6 组件速查表
#组件核心数据结构手写 vs 框架本仓对应
①State Managermessages[]手写就 1 个 listexperiment_modules/solving/agent_runtime/ 各 adapter 的 message 维护
②Tool Registrydict[str, callable]框架是装饰器 @tool各 brand adapter 的 tools/ 子目录
③Output Parser解析后的结构体框架是 Pydantic modelbase_one_shot_runner.py 的 _extract_diff()
④Loop Controllerwhile/for + 3 守门框架是 Graph 节点编排stages/s4/worker_entry.py 的 walltime_monitor
⑤Error Handlertry/except + 重试框架是 retry policystages/s4/ 的 subprocess 退出码协议(0/1/124)
⑥LoggerJSONL 文件框架是 callback hooks4_worker.log + result.json + experiments.db
spec §3 harness 拆件 ref: agent_runtime/registry

第四讲 · LLM 在 loop 里:知道什么 / 不知道什么

第四讲 · 划清 LLM 的"知识边界"(1 对照表 + 3 个常见误解)
为什么 LLM 会"幻觉"?为什么 context window 重要?答案都在 LLM 不知道什么里。
第 4 讲 · 1 对照表 + 3 个误解

LLM 知道的 vs 不知道的:每次调用是个"无状态函数"

LLM 每次被调用都是一个无状态函数:给它 messages + tools,它返回 content + tool_calls。它不会"记住"上一轮,不会"知道"你之前跑过什么工具、用了多少 token、剩多少预算。所有这些信息都得由 harness 显式塞进 messages 里,LLM 才能"看见"。

LLM 知道什么 / 不知道什么(每次调用层面)
维度LLM 知道LLM 不知道(除非 harness 显式塞)
对话历史✓ messages[] 里塞的全部内容历史中被摘要掉的早期轮次
工具定义✓ tools 参数里声明的 schema工具实际执行的副作用(如写了哪个文件)
工具结果✓ role:"tool" 消息里的 content工具抛错时的堆栈(除非 harness 转成文本)
当前时间✗(除非 system 注入)真实的"现在";模型内置知识有截止日期
自己的身份✗(除非 system 注入)自己正在被哪个 harness 调用、第几轮
token 用量✗已用多少、剩多少、会不会爆 context window
任务进度✗"我在第 3 步 / 总共 10 步" / "目标完成度"
3 个由"不知道"引发的常见误解
误解 1:"LLM 怎么会犯这么简单的错?它不是看过我之前问的吗?"
真相:它看过的是harness 当时塞进 messages 的版本。如果你的 State Manager 在第 5 轮把前 4 轮摘要成一句话,messages[0] 就不是原始问题了。早期细节丢失 ≠ LLM 健忘,是harness 上下文管理的锅。
误解 2:"为什么 LLM 调工具时报错后就不再重试了?"
真相:LLM 根本不知道自己上一次调工具失败——除非 harness 把错误信息转成 role:"tool" 消息里的文本塞回去("工具执行失败:FileNotFoundError")。错误可见性 = harness 的责任。
误解 3:"为什么 LLM 在第 20 轮突然开始胡言乱语?"
真相:Context window 满了,harness 没做摘要压缩,新消息覆盖了早期消息,LLM 看到的"历史"已经残缺。这是Loop Controller 的第 3 个守门(max_tokens)和 State Manager 的摘要策略该处理的事。
spec §4 LLM 边界 ref: OpenAI Function Calling

第五讲 · The Loop:主循环伪代码 + ownership 注释

第五讲 · 把所有组件拼成一段循环(1 伪代码 + 1 时序图)
用一段 30 行的伪代码 + 一张时序图,把前 4 讲的抽象全收回来。
第 5 讲 · 1 伪代码 + 1 时序图

主循环伪代码:每行注释 ownership(harness / LLM / tool)

把 6 大组件、3 大守门、所有"知道/不知道"的边界,收成一段可读的伪代码。读完这段你就拿到了 Agent 开发的"主模板"——LangGraph、AutoGen、各 brand adapter 的源码都是这个模板的变体。

主循环伪代码(每行 # 注释标了 owner)
def run_agent(user_query: str, max_iter: int = 20, wall_time: int = 600):
    # ====== ① 初始化:State Manager ======
    messages = [{"role": "user", "content": user_query}]              # [harness]
    tools = declare_tools()                                          # [harness] 工具定义
    start = time.time()                                              # [harness] Loop 守门 2 的起点

    # ====== ② 主循环:Loop Controller ======
    for turn in range(max_iter):                                     # [harness] 守门 1:轮次
        if time.time() - start > wall_time:                          # [harness] 守门 2:墙钟
            return TimeoutError(f"超出 {wall_time}s")

        # ====== ③ 调 LLM ======
        response = client.chat.completions.create(                   # [LLM] ←—— 唯一的"AI 时刻"
            model=MODEL, messages=messages, tools=tools,
            max_tokens=4096,                                         # [harness] 守门 3:单次 token 上限
        )
        log(response.usage)                                          # [harness] Logger

        # ====== ④ Output Parser ======
        msg = response.choices[0].message                            # [harness] 拆出 assistant 消息
        messages.append(msg)                                         # [harness] State: 写回
        log({"turn": turn, "msg": msg})                              # [harness] Logger

        # ====== ⑤ 分支 ======
        if msg.finish_reason == "stop":                              # [harness] 自然结束
            return msg.content                                       # → 返回最终答案

        if msg.finish_reason == "length":                            # [harness] 截断
            raise LengthError("撞 max_tokens,需扩预算或续写")

        if msg.tool_calls:                                           # [harness] 模型要调工具
            for tc in msg.tool_calls:                                # [harness] 多个并行调用
                args = json.loads(tc.function.arguments)             # [harness] 字符串 → 对象
                try:
                    result = TOOL_REGISTRY[tc.function.name](**args) # [tool]  ←—— 副作用在这里
                except Exception as e:
                    result = f"工具执行失败:{e}"                     # [harness] Error Handler 兜底
                messages.append({                                    # [harness] State: 写回 tool 结果
                    "role": "tool", "tool_call_id": tc.id,
                    "content": str(result),
                })
            continue                                                 # [harness] 回到循环顶,再调一次 LLM

    raise MaxIterError(f"超出 {max_iter} 轮,可能死循环")            # [harness] 守门 1 触发
图 2 · 主循环时序图(蓝色 harness / 橙色 LLM / 绿色 tool)
sequenceDiagram autonumber actor U as 用户 participant H as 🟦 Harness participant L as 🟧 LLM participant T as 🟩 Tool U->>H: 提交 query loop 最多 N 轮 / M 秒 H->>L: messages + tools + max_tokens L-->>H: {content, tool_calls?, finish_reason} alt finish_reason = "stop" H-->>U: 返回 final answer else finish_reason = "tool_calls" loop 每个 tool_call H->>T: 执行 tool(arguments) T-->>H: result(或 exception) H->>H: 错误兜底 H->>H: messages.append({role:"tool", tool_call_id, content}) end H->>H: 回到循环顶 else finish_reason = "length" H-->>U: 报错 LengthError end end
图 2 · 一次完整 Agent 循环的时序:Harness(蓝)每轮都跑、LLM(橙)只在第 4 步被调、Tool(绿)只在 tool_calls 分支跑
看图要点:① 整个循环里 LLM 只在第 4 步被调一次——其他全是 harness 在跑;② 状态续接靠 harness 的 messages.append(LLM 自己不会记);③ 错误兜底是 harness 的 try/except(LLM 不知道工具抛错);④ 三种退出路径:stop(自然)/ length(截断)/ MaxIterError(守门 1)。
spec §5 主循环模板 paper: ReAct (ICLR 2023)

第六讲 · 真实例子:SWE-bench S1-S7 pipeline 怎么对应到 harness

第六讲 · 工业级例子(1 映射表 + 1 result.json 节选 + 2 失败案例)
把抽象落地:本仓 S1-S7 编排器是 harness loop 的工业级实现,逐 stage 标注对应到前 4 讲哪个组件。
第 6 讲 · 1 映射 + 1 真实数据 + 2 案例

本仓 S1-S7:6 个 stage 怎么对应到 harness loop

本仓 experiment_modules/orchestration/ 有一个完整的工业级 Agent pipeline:出题 → 准备环境 → 作答 → 抽 diff → 评分 → 记录。看着是 6 个 stage,核心循环只发生在 S4_solve 一个 stage 里——其他 5 个都是它的"周边"。下表把映射关系摊开。

S1-S7 stage × harness 6 组件 映射表
Stage职责对应 harness 组件平均时长
S1_build出题(拿 instance_id + 测试用例)State Manager 初始化~0.1s
S2_prepare拉 Docker 镜像(amd64 + Rosetta)Loop 之前的 prep(不算 loop)~17s
S4_solve跑 agentic 循环(这才是 Agent)完整 harness loop(卡片 5 的伪代码)~420s
S5_patch从 trajectory 抽 model_patchOutput Parser 的 _extract_diff()~0.01s
S6_score在容器里跑 F2P(Fail-to-Pass,修复前失败、修复后应通过的测试)/ P2P(Pass-to-Pass,修复前后都应通过的回归测试)测试Loop 之后的独立 verifier~160s
S7_record写 result.json + experiments.dbLogger 持久化部分~0.1s
真实 result.json 节选(一次 django 任务)
{
  "instance_id": "django__django-10914",
  "model": "qwen-code",
  "resolved": true,                            // ← S6 评分结果:F2P + P2P 全过
  "stages": {
    "S1_build": "done", "S2_prepare": "done",
    "S4_solve": "done", "S5_patch": "done",
    "S6_score": "done", "S7_record": "done"
  },
  "stage_timings": {                            // ← Loop Controller 的 3 大守门日志
    "S1_build": 0.1, "S2_prepare": 16.8,
    "S4_solve": 420.4,                          // ← 420s = 7 分钟 agentic 循环
    "S5_patch": 0.0, "S6_score": 162.4
  },
  "adapter_attempts": [{
    "adapter": "qwen-agent", "subprocess": true,
    "wall_time_seconds": 420.3, "returncode": -15,
    "killed": true, "rescued": true,            // ← 被 SIGTERM(墙钟超时),但已被 S5 抽到 diff
    "outcome": "resolved"
  }]
}
2 个真实失败案例(理解 harness 的"边界")
案例 1 · tool_call_id 配对错误(Output Parser 失职)
现象:S4 跑到第 8 轮突然 400,循环中断。
原因:harness 把 tool 消息的 tool_call_id 写错(截断了前缀),OpenAI 报 "messages with role 'tool' must be a response to a preceeding message with 'tool_calls'"。
修复:Error Handler 加一层 assert msg.tool_calls[i].id == tool_msg.tool_call_id,错了立刻抛错而不是继续跑。
案例 2 · 墙钟超时(Loop Controller 兜底)
现象:上面 result.json 的 returncode: -15 + killed: true——子进程被 SIGTERM。
含义:harness 跑满 420s 墙钟上限(守门 2 触发),主进程发 SIGTERM 杀子进程;S5 抽到部分 diff,S6 仍然能跑(rescued: true)。
设计启示:守门 2 不是"失败",是"止损"——能拿部分结果就别全扔。
ref: orchestration/README.md ref: agent_runtime/registry

第七讲 · 最小可运行 harness + 4 种主流模式

第七讲 · 30 行 harness + 4 模式对比 + 选型心法(2 代码 + 1 表 + 1 心法)
理论收尾,给一份能跑的代码 + 一张选型表。
第 7 讲 · 2 代码 + 1 模式表 + 1 选型心法

30 行最小可运行 harness(拷走就能跑)

用 OpenAI 官方 SDK 写一个最小 harness:模型可以调 get_weather 工具查天气,多轮循环直到 finish_reason="stop"。整个文件 30 行,对应卡片 5 的伪代码每一行。

最小 harness 完整代码(30 行 · 可运行)
import json, openai
client = openai.OpenAI()                                          # 从 OPENAI_API_KEY 读

def get_weather(city: str) -> str:                                 # 工具实现
    return f"{city} 25°C 晴"

messages = [{"role": "user", "content": "北京今天几度?"}]
tools = [{"type": "function", "function": {
    "name": "get_weather", "description": "查天气",
    "parameters": {"type": "object",
                   "properties": {"city": {"type": "string"}},
                   "required": ["city"]}}}]

for turn in range(10):                                            # 守门 1
    r = client.chat.completions.create(
        model="gpt-4o-mini", messages=messages, tools=tools)
    msg = r.choices[0].message
    messages.append(msg)                                           # State
    if msg.tool_calls:                                             # 工具分支
        for tc in msg.tool_calls:
            args = json.loads(tc.function.arguments)
            messages.append({"role": "tool",
                             "tool_call_id": tc.id,
                             "content": get_weather(**args)})
    else:                                                          # 自然结束
        print(msg.content)
        break
对照卡片 5:上面 30 行 = 卡片 5 伪代码的Python 化。每个 # 守门 N / # State / # 工具分支 注释都对应得上。读懂卡片 5 → 就能读懂所有 Agent 框架的源码。
4 种主流 agent 模式对比

主循环的形状决定了 Agent 的"性格"。同样 6 大组件,循环形状不同就变成不同的范式。

模式循环形状适合失败模式
ReAct
(Reason+Act)
每轮:think → act → observe,
循环到模型说停
通用任务 / 探索 / debug长链推理时容易"绕";
无显式计划
Plan-and-Execute第 1 轮:模型出完整计划 →
后续轮:逐项执行
可分解的多步任务 / 工具编排计划错误导致全盘错;
计划改不了
Reflection执行 → 自我批判 →
重做 → 再批判
代码生成 / 写作 /
质量敏感场景
批判无依据陷入死循环;
token 翻倍
Multi-Agent多个 harness 协作(planner /
executor / critic)
角色分工 / 复杂决策 /
需要不同视角
通信开销;
角色边界模糊
选型心法(4 个问题定模式)
  1. 任务能"一眼看穿"步骤吗? 能 → Plan-and-Execute;不能 → ReAct
  2. 结果对质量敏感吗(代码 / 文档 / 数据)? 是 → Reflection;否 → ReAct 即可
  3. 需要多个"角色"协作吗(计划 / 执行 / 评审)? 是 → Multi-Agent;否 → 单 loop 起步
  4. 成本敏感吗? 敏感 → 优先 ReAct(每轮 token 最少);不敏感 → Reflection / Multi-Agent 都可以试
个人心法(从工程实践角度):从 ReAct 起步,跑通最小环(卡片 7 的 30 行)→ 加工具(卡片 3 的 Tool Registry)→ 跑不通的时候再升级到 Plan-and-Execute / Reflection。Multi-Agent 是最后的手段,不是起点——多个 harness 出 bug 时排查链路指数膨胀。
spec §7 模式与选型 paper: ReAct (ICLR 2023) paper: Reflexion (NeurIPS 2023)

收口:把 7 张卡片装回一个心智

一页总结:Agent 不是一个新物种,是一种编程范式。Agent = harness(你写的循环代码)+ LLM(外部 API)。
  • 卡片 1:LLM API = 单回合函数,不"记住"前文
  • 卡片 2:Agent = harness + LLM + 外部世界;ownership 边界一清二楚
  • 卡片 3:harness 6 大组件:State / Tools / Parser / Loop / Error / Logger
  • 卡片 4:LLM 只知道 messages 里的内容;其他全靠 harness 显式塞
  • 卡片 5:主循环伪代码 + 时序图 = 30 行可执行模板
  • 卡片 6:本仓 S1-S7 = harness loop 的工业级实现,S4_solve 才是 Agent 本体
  • 卡片 7:30 行可运行 harness + 4 种模式选型心法

读完后,你下次再看到"AI Agent 自动完成了 XX 任务"——脑子里浮现的不再是"魔法",而是:一段循环代码(harness)在反复调大模型(LLM),中间穿插工具调用和状态管理。这就是"理解 Agent 开发"的核心:把"AI 智能"还原成"代码逻辑"。