> Agent Harness Patterns 第 12 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/plan-and-reflect · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 先规划，再反思

动手之前先把计划写下来，接受答案之前先检查工作。计划让智能体不偏离方向；检查能抓住它声称做了却没做的事。

## 问题

给智能体一份由好几部分组成的工作，它会从眼前有什么就先做什么。做到一半，最初的几步已经落在历史记录很靠前的位置，它会忘掉其中一步。最后它信心满满地回答“搞定！”，因为没有任何东西逼它去看一眼。

这就是 Scatterbrain。两种失败合二为一：没有计划，所以步骤会丢；没有检查，所以出了错的步骤被报告为已完成。在探戈街，工具明明白白地说报纸掉进了灌木丛。模型读到了，却照样勾上了格子。

## 解决方案

**先规划。**在第一个真正的动作之前，智能体把工作写成一份任务清单，每完成一项就勾掉一项。让它奏效的诀窍是：整份计划会被加进每一次请求，所以不管历史记录变得多长，模型总能看到哪些做完了、哪些还剩着。在 astorlm 里，这就是 `pattern: 'PLAN_EXECUTE'`：两个工具 `add_plan_item` 和 `update_plan_item`，外加每一轮都放进 system prompt 的计划。

**接受之前先反思。**一个打了勾的格子，只是模型的说法。当 `run()` 返回时，你的代码在接受结果之前先检查它。如果缺了什么，这个发现会作为一条新消息送回*同一个*会话：智能体保留它的历史记录和计划，只修正出错的部分。要限制轮次。检查有三种方式，从强到弱：

- **检查现实世界**
   你的代码查看结果本身：门廊、数据库里的行、测试套件、磁盘上的文件。
   只要能用上，它就是最好的检查者：便宜、精确，而且谁也说服不了它。前提是工作能被代码验证。
- **一个评审模型**
   第二次调用阅读任务、答案和一份检查清单，列出哪里错了、缺了什么。
   适合没有代码能检查的工作：一段总结、一封邮件、一份计划。要多花一次调用，它也可能漏看，而且需要具体的标准，而不是“这个好不好？”。
- **问智能体自己**
   system prompt 要求智能体在回答之前重读一遍自己的工作。
   免费，有时也够用。但这是同一个模型在给自己打分，带着同样的盲点：它已经勾过一次 14 号了。

这和第 11 关的评估是同一个思路，只不过用在运行时：评估在事后给运行打分，用来改进智能体；检查则在用户看到之前，给这一次运行打分。

## 角色

还是那群熟悉的角色，这次换成了一趟送报。

- **报亭** (模型): 柜台后面的神谕者。它决定每一步，却从不骑车。
- **路线单** (计划): 任务和它们的格子。每一轮它都会闪一下金光：它随每次请求一起发送。
- **Astor 的自行车** (循环): 把每个工具调用送出去再带回来，历史记录装在他的手风琴里。
- **一次投递** (deliver): 一个普通工具。它会说报纸落在了哪里。
- **编辑** (你的代码): 交代任务，并在接受答案之前检查各家门廊。它的放大镜就是 `review()`。

## 代码

**使用 astorlm：**`pattern: 'PLAN_EXECUTE'` 会加上计划工具，并把计划放进每次请求；`getPlan()` 可以把它读回来。检查就是 `run()` 之后的普通代码，而在同一个智能体上再调用一次 `run()`，会延续同一个会话。

**从零手写：**一个列表、两个修改它的工具，以及一个每轮都带着这个列表重新拼出来的 system prompt。历史记录放在 `run()` 外面，所以修正会延续同一段对话。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

const SUBSCRIBERS = [12, 14, 18]
const porches = new Set<number>() // the real world: which porches have a paper

const deliver = tool({
  name: 'deliver',
  description: 'Ride to a house and throw today’s paper onto its porch. Says where the paper landed.',
  schema: z.object({ house: z.number() }),
  execute: async ({ house }) => {
    const landed = throwPaper(house) // your code: 'porch' or 'bushes'
    if (landed === 'porch') porches.add(house)
    return landed === 'porch' ? `Paper on the porch at #${house}.` : `Paper landed in the bushes at #${house}.`
  },
})

// PLAN: the agent gets add_plan_item and update_plan_item,
// and the current plan is added to the system prompt on every turn.
const agent = await createLocalAgent({
  provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
  pattern: 'PLAN_EXECUTE',
  systemPrompt: 'You deliver newspapers. Plan every stop before you start, and tick each task as you go.',
  tools: [deliver],
  maxTurns: 20,
})

// REFLECT: check the work itself before accepting the answer. Deterministic when you can;
// a second model with a rubric when you can't.
const review = (): string[] => SUBSCRIBERS.filter((house) => !porches.has(house)).map((house) => `#${house} has no paper on the porch`)

let answer = await agent.run(`Deliver today’s paper to every subscriber on Tango Street: ${SUBSCRIBERS.join(', ')}.`)
for (let round = 1; round <= 2; round++) {
  const problems = review()
  if (problems.length === 0) break
  // Same agent, same session: it keeps its history and its plan, and fixes what's missing.
  answer = await agent.run(`Review found: ${problems.join('; ')}. Fix it.`)
}

console.log(agent.getPlan()) // [{ id: '1', description: 'Deliver to #12', status: 'completed' }, …]
console.log(answer.content)
```

**TypeScript**

```ts
// Plan and reflect, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

const SUBSCRIBERS = [12, 14, 18]
const porches = new Set<number>() // the real world: which porches have a paper

// 1. The plan: a list the model writes and ticks with two tools.
type Task = { id: number; description: string; status: 'pending' | 'completed' }
const plan: Task[] = []

type ToolFn = (args: Record<string, string | number>) => string
const tools: Record<string, ToolFn> = {
  add_plan_item: ({ description }) => {
    plan.push({ id: plan.length + 1, description: String(description), status: 'pending' })
    return `Task added with ID: ${plan.length}`
  },
  update_plan_item: ({ id, status }) => {
    const task = plan.find((t) => t.id === Number(id))
    if (!task) return `No task ${id}`
    task.status = status === 'completed' ? 'completed' : 'pending'
    return `Task ${id} status updated to ${task.status}.`
  },
  deliver: ({ house }) => {
    const landed = throwPaper(Number(house)) // your code: 'porch' or 'bushes'
    if (landed === 'porch') porches.add(Number(house))
    return landed === 'porch' ? `Paper on the porch at #${house}.` : `Paper landed in the bushes at #${house}.`
  },
}
const toolSchemas = [/* one JSON Schema per tool: add_plan_item(description), update_plan_item(id, status), deliver(house) */]

// 2. The plan goes into the system prompt on EVERY turn, so the model never loses track.
const system = (): string =>
  'You deliver newspapers. Plan every stop with add_plan_item before you start, and tick each task as you go.\n' +
  (plan.length ? plan.map((t) => `- [${t.status}] ${t.description} (id ${t.id})`).join('\n') : '(no plan yet)')

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// The history lives outside run(), so a second run continues the same session.
const messages: Message[] = []

async function run(prompt: string, maxTurns = 20): Promise<string> {
  messages.push({ role: 'user', content: prompt })
  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: system() }, ...messages], tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const fn = tools[call.function.name]
      const output = fn ? fn(JSON.parse(call.function.arguments)) : `Unknown tool: ${call.function.name}`
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// 3. Reflect: check the work itself before accepting the answer, and send what's missing back.
const review = (): string[] => SUBSCRIBERS.filter((house) => !porches.has(house)).map((house) => `#${house} has no paper on the porch`)

let answer = await run(`Deliver today’s paper to every subscriber on Tango Street: ${SUBSCRIBERS.join(', ')}.`)
for (let round = 1; round <= 2; round++) {
  const problems = review()
  if (problems.length === 0) break
  answer = await run(`Review found: ${problems.join('; ')}. Fix it.`)
}
console.log(answer)
```

**Python**

```python
# Plan and reflect, from scratch. Standard library only, no SDK.
import json
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

def post(path, payload):
    request = urllib.request.Request(
        f"{LLM['base_url']}{path}",
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

SUBSCRIBERS = [12, 14, 18]
PORCHES = set()  # the real world: which porches have a paper

# 1. The plan: a list the model writes and ticks with two tools.
PLAN = []

def add_plan_item(description):
    PLAN.append({"id": len(PLAN) + 1, "description": description, "status": "pending"})
    return f"Task added with ID: {len(PLAN)}"

def update_plan_item(id, status):
    task = next((t for t in PLAN if t["id"] == int(id)), None)
    if task is None:
        return f"No task {id}"
    task["status"] = "completed" if status == "completed" else "pending"
    return f"Task {id} status updated to {task['status']}."

def deliver(house):
    landed = throw_paper(house)  # your code: "porch" or "bushes"
    if landed == "porch":
        PORCHES.add(house)
        return f"Paper on the porch at #{house}."
    return f"Paper landed in the bushes at #{house}."

TOOLS = {"add_plan_item": add_plan_item, "update_plan_item": update_plan_item, "deliver": deliver}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool: add_plan_item(description), update_plan_item(id, status), deliver(house)

# 2. The plan goes into the system prompt on EVERY turn, so the model never loses track.
def system():
    lines = [f"- [{t['status']}] {t['description']} (id {t['id']})" for t in PLAN] or ["(no plan yet)"]
    return "You deliver newspapers. Plan every stop with add_plan_item before you start, and tick each task as you go.\n" + "\n".join(lines)

# The history lives outside run(), so a second run continues the same session.
MESSAGES = []

def run(prompt, max_turns=20):
    MESSAGES.append({"role": "user", "content": prompt})
    for _ in range(max_turns):
        request = [{"role": "system", "content": system()}, *MESSAGES]
        choice = post("/chat/completions", {"model": LLM["model"], "messages": request, "tools": TOOL_SCHEMAS})["choices"][0]
        reply = choice["message"]
        MESSAGES.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
            MESSAGES.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# 3. Reflect: check the work itself before accepting the answer, and send what's missing back.
def review():
    return [f"#{house} has no paper on the porch" for house in SUBSCRIBERS if house not in PORCHES]

answer = run(f"Deliver today's paper to every subscriber on Tango Street: {', '.join(map(str, SUBSCRIBERS))}.")
for _ in range(2):
    problems = review()
    if not problems:
        break
    answer = run(f"Review found: {'; '.join(problems)}. Fix it.")
print(answer)
```

## 注意事项

- **任务要小，而且要能验证。**“送到 14 号”可以验证；“搞定这条街”不行。你能验证的任务，检查才抓得住。
- **允许计划改变。**计划总会撞上现实：一条街封了，一位顾客取消了。智能体应该能增加、删除或重排任务，而不是死守一份过时的清单。
- **检查现实世界，而不是计划。**单子上是三个勾。只检查单子的话就通过了。检查必须去看结果本身。
- **限制轮次。**一个永远通不过的检查，或者一个修不好自己发现的问题的智能体，会永远循环下去。两到三轮，然后交给人来处理（第 13 关）。
- **一句话就能做完的事，别做计划。**规划要花轮次和 token。只是查一次东西的话，就跳过它。当工作有好几步、容易漏掉时，它才划算。

## 相关模式

- [2 · 智能体循环](https://harnesspatterns.dev/zh/patterns/agent-loop.md)
- [11 · 可观测性与评估](https://harnesspatterns.dev/zh/patterns/observability.md)
- [10 · 每圈重新开始](https://harnesspatterns.dev/zh/patterns/fresh-laps.md)
- [13 · 人在回路](https://harnesspatterns.dev/zh/patterns/human-in-the-loop.md)
- [15 · 子智能体](https://harnesspatterns.dev/zh/patterns/subagents.md)
