> Agent Harness Patterns 第 2 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/agent-loop · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 智能体循环

单靠模型本身什么也做不了：它只会生成文本。智能体循环才是把它变成智能体的东西。循环把历史记录交给模型，执行模型申请的工具，把结果交回去，然后一直重复，直到模型说“完成”。

## 问题

鸟妈妈问模型：“你能把猪的堡垒清掉吗？”模型什么都发射不了。它最多只能回复一个工具*请求*：`{ name: "launch_red", input: { angle: 40 } }`。

如果你的代码只调用一次模型，对话就到此为止了。你手里只剩一个没人执行的请求，没有答案。如果你手动执行这个工具，下一轮又会碰到同样的问题，因为模型可能还需要另一个工具，然后再一个。

## 解决方案

一个只有一条退出规则的循环：

1. 把**完整的历史记录**加上可用工具列表发给模型。
2. 如果回复以 `stopReason: "end_turn"` 结束，就返回它。这是唯一的正常出口。
3. 如果以 `"tool_use"` 结束，就执行每个被申请的工具，把结果以 `tool_result` 块的形式追加到历史记录里，然后回到第 1 步。

模型决定*做什么*；循环负责*去做*。这种分工是其他所有模式的基础：其余的一切（steering、上下文压缩、子智能体）都挂在这个循环的某个点上。

## 角色

把循环讲成一段小小的冒险。认识了这些角色之后，就没什么需要再破解的了。

- **Astor** (循环): 一个小小的探戈舞者，也是唯一会走动的角色。他把问题带给神谕者，每一发都跑到弹弓前，再把答案带回给鸟妈妈。
- **神谕者** (Provider): 模型。它从不碰弹弓：只听手风琴，然后递回一张纸条。需要工具时是橙色，完成时是金色。
- **班多钮手风琴** (messages[]): 历史记录，每条消息是一褶彩色的风箱。风箱每一圈都在变长，而神谕者每次都要把每一褶都听一遍。那些往神谕者飘去的音符就是这个意思。
- **长椅** (ToolRegistry): 每个工具一只鸟：`launch_red`、`launch_bomb`，还有一只今天没人需要。Astor 发射纸条上写的那只，再举起结果：成功就是绿色。
- **鸟妈妈** (agent.run()): 你的代码。它提出问题，然后等待。
- **轨迹、分数和小鸟** (历史记录、tokens、maxTurns): 每一发都在天空留下轨迹，就像历史记录保留每个结果一样。分数就是 token，而且因为整个历史记录都要重新发送，每一圈都涨得更多。每一圈都要从顶栏那排小鸟里扣掉一只，那就是轮数预算。数字仅作示意。

EventBus 面板展示真实循环在动画每一步发出的事件。

## 代码

**使用 astorlm：**同一个循环就在 `src/agent/loop.ts` 里，带有流式输出、重试、钩子、并行执行工具和取消功能。从外面看，它是这样的。

**从零手写：**大约 40 行代码，可对接任何兼容 OpenAI 的端点，不用 SDK：TypeScript 里只用 `fetch`，Python 里只用标准库。上面的三个步骤都在注释里标出来了。把开头的 `LLM` 配置块换成你自己的端点、模型和密钥。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

// Your game's functions, wrapped as tools: one per bird.
const angle = z.number().min(10).max(80).describe('Launch angle in degrees')

const launchRed = tool({
  name: 'launch_red',
  description: 'Fling the red bird. Good against wood. Returns what fell and how many pigs are left.',
  schema: z.object({ angle }),
  execute: async ({ angle }) => level.fling('red', angle), // your code
})

const launchBomb = tool({
  name: 'launch_bomb',
  description: 'Fling the bomb bird. It explodes on impact: the one to use against stone.',
  schema: z.object({ angle }),
  execute: async ({ angle }) => level.fling('bomb', angle), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [launchRed, launchBomb],
  maxTurns: 10, // the birds in line: a cap on the laps
})

agent.on('tool-start', (tool) => console.log('→', tool.name, tool.input))

const answer = await agent.run('The pigs took our eggs! Can you clear their fort?')
```

**TypeScript**

```ts
// Agent loop from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { launch_red: launchRed, launch_bomb: launchBomb }
const toolSchemas = [/* one JSON Schema per tool */]

export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]

  for (let turn = 1; turn <= maxTurns; turn++) {
    // 1. Send the whole history plus the tool list.
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)

    // 2. No tool calls: the model is done. The only normal exit.
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    // 3. Run each requested tool and feed the result back as a message.
    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      if (run) {
        try {
          output = await run(JSON.parse(call.function.arguments))
        } catch (err) {
          output = `Error: ${err instanceof Error ? err.message : err}`
        }
      }
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
```

**Python**

```python
# Agent loop from scratch. Standard library only, no SDK.
import json
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"launch_red": launch_red, "launch_bomb": launch_bomb}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

def run_agent(prompt, max_turns=10):
    messages = [{"role": "user", "content": prompt}]

    for _ in range(max_turns):
        # 1. Send the whole history plus the tool list.
        choice = chat(messages)
        reply = choice["message"]
        messages.append(reply)

        # 2. No tool calls: the model is done. The only normal exit.
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        # 3. Run each requested tool and feed the result back as a message.
        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            run = TOOLS.get(name)
            try:
                output = run(**json.loads(call["function"]["arguments"])) if run else f"Unknown tool: {name}"
            except Exception as err:
                output = f"Error: {err}"
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")
```

注意，工具出错并不会让循环停下：错误会以文本的形式回到模型那里，让它在下一轮自我纠正。

## 何时使用，以及要注意什么

**只要模型需要在回答之前先行动**（读取、搜索、运行某些东西），就用它。如果你只需要转换文本，一次调用就够了，而且更便宜。

- **成本每一圈都在增长。**留意手风琴和 token 计数器：神谕者每一轮都要把每一褶都听一遍，所以每一圈都会把之前的一切重新发送一次。轮数一多，上下文就会被填满。
- **一定要设置 `maxTurns`。**一个犯糊涂的模型可能会永远不停地申请工具。没有上限，循环也会永远跑下去。

> 真实故障
>
> 在一次多轮运行中，代理（proxy）把请求故障转移到了一个更弱的免费模型上。这个模型陷入了毫无意义的推理循环，始终没有返回 `end_turn`。最后是一个 60 秒的超时把它停下的，而不是循环本身。`maxTurns` 能防止轮数太多，却防不住一轮永远不结束：为此你需要一个超时或者看门狗（在没有进展时切断运行的机制）。

## 相关模式

- [4 · 何时停止](https://harnesspatterns.dev/zh/patterns/when-to-stop.md)
- [5 · 循环中的错误](https://harnesspatterns.dev/zh/patterns/errors-in-the-loop.md)
- [6 · 钩子](https://harnesspatterns.dev/zh/patterns/hooks.md)
- [10 · 每圈重新开始](https://harnesspatterns.dev/zh/patterns/fresh-laps.md)
