> Agent Harness Patterns 第 5 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/errors-in-the-loop · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 循环中的错误

工具会出错。模型会传来错误的输入。服务器会宕机一秒。对每一种失败，智能体循环都要决定由谁来处理：模型、你的代码，还是谁都不管。

## 问题

一个智能体在玩冒险游戏。它的工具是屏幕底部的动词 `look_at` 和 `use`，任务是“打开神殿的门”。模型做了最显而易见的事：用钥匙开门。可是锁生锈了，工具抛出了异常。

面对这种时刻，老式冒险游戏分成两派。有的游戏里，走错一步你就死了：game over，回到上一个存档。另一些游戏里你死不了；游戏只会告诉你为什么行不通，然后你换个办法再试。智能体循环也必须选一派。

如果循环不接住这个错误，Krash 就赢了：异常沿着循环一路往上冒，`run()` 抛出异常。整次运行就这么丢了，而错误里明明写着该怎么做：“先上油”。唯一能用上这条信息的读者，偏偏没收到。接住它，作为工具结果送回去，模型读到后就会再试一次。这只花一轮。不接住，代价是整次运行。

## 三种失败

- 模型申请错了
   输入无效、JSON 损坏、调用了不存在的工具、钥匙转不动。
   **交给模型。** 错误以标记为错误的工具结果返回。模型读到后，下一轮换个办法试。
- 模型的服务器出了故障
   429（请求太多）、5xx（服务器问题）、连接中断。
   **交给你的代码，重试。** 等一会儿再调用，每次都多等一点。模型永远不会知道：根本没有产生需要它读的轮次。
- 谁也修不好
   401（API 密钥错误）、你代码里的 bug、用户取消了。
   **谁都不管。** 让它直接抛出。重试没有用，把它瞒着模型，只会让模型去瞎猜。

诀窍在于别把它们混在一起。重试一个 400，只会把同一个坏掉的请求再发一遍。把 503 拿给模型看，只会浪费一轮在它修不了的东西上。而接住你自己代码里的 bug，只会把它藏起来。

> astorlm 现状
>
> 工具错误已经替你处理好了：ToolRegistry 会接住未知工具、没通过 Zod schema 的输入，以及 `execute` 抛出的任何东西，并以 `is_error: true` 送回去。对模型的重试**默认是关闭的**：不设置 `retry` 选项的话，一个 503 就会让运行以 `session_end: error` 结束。另外，schema 错误是以原始的 ZodError 交给模型的，也就是把每个问题都倒成一大段 JSON。能用，但换成你自己写的一句简短说明会更好读。

## 代码

**使用 astorlm：**工具只管抛出异常，接住的活交给 astorlm。你为模型的服务器打开 `retry`，再监听 EventBus，就能看到两层机制各自在工作。

**从零手写：**第 2 关的循环，每一层对应一个函数：`callModel` 带退避地重试服务器，`runTool` 把每次工具失败都变成模型读得懂的文本，其余的一律放行抛出。在此之上，还有一个小计数器：当模型一直撞同一堵墙时，就放弃。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const room = { lockOiled: false, doorOpen: false }

const use = tool({
  name: 'use',
  description: 'Use an inventory item with something in the room.',
  schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
  execute: async ({ item, target }) => {
    if (item === 'oil_can' && target === 'lock') {
      room.lockOiled = true
      return 'You oil the lock. It looks like it might turn now.'
    }
    if (item === 'key' && target === 'door') {
      // Just throw. astorlm catches it and sends the message back to the model, marked is_error.
      if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
      room.doorOpen = true
      return 'Click! The key turns and the door swings open.'
    }
    throw new Error(`Nothing happens when you use ${item} with ${target}.`)
  },
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [lookAt, use],
  // Off by default: without it, a single 429 ends the run.
  retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
  maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})

// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
  if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
  if (event.type === 'provider_retry') {
    console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
  }
})

try {
  const last = await agent.run('Open the temple door.')
  console.log(last.content)
} catch (err) {
  // Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
  console.error('The run failed:', err)
}
```

**TypeScript**

```ts
// Errors in an agent loop, from scratch. Plain fetch, no SDK.
// The agent plays an adventure game: its tools are the verbs LOOK AT and USE.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

const room = { lockOiled: false, doorOpen: false }

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { look_at: lookAt, use }
const toolSchemas = [/* one JSON Schema per tool */]

async function use({ item, target }: Record<string, string>): Promise<string> {
  if (item === 'oil_can' && target === 'lock') {
    room.lockOiled = true
    return 'You oil the lock. It looks like it might turn now.'
  }
  if (item === 'key' && target === 'door') {
    // Say what went wrong and what to try instead: this text is all the model will get.
    if (!room.lockOiled) throw new Error('the lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
    room.doorOpen = true
    return 'Click! The key turns and the door swings open.'
  }
  throw new Error(`nothing happens when you use ${item} with ${target}.`)
}

// Layer 1, the model's server. Retry only what is likely to pass by itself.
async function callModel(messages: Message[], maxAttempts = 3) {
  for (let attempt = 1; ; attempt++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    }).catch(() => null) // null: no response at all, the network failed
    if (res?.ok) return (await res.json()).choices[0]

    // 429 (too many requests), 5xx (server trouble) and network errors tend to pass.
    // A 400 or a 401 won't: retrying just sends the same broken request again.
    const transient = res === null || res.status === 429 || res.status >= 500
    if (!transient || attempt === maxAttempts) throw new Error(`Model call failed: ${res?.status ?? 'network error'}`)
    // Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
    await new Promise((resolve) => setTimeout(resolve, Math.random() * 1000 * 2 ** (attempt - 1)))
  }
}

// Layer 2, the tools. Whatever goes wrong becomes a result the model can read.
async function runTool(call: ToolCall): Promise<{ output: string; isError: boolean }> {
  const run = tools[call.function.name]
  if (!run) return { output: `Unknown tool: ${call.function.name}. Tools: ${Object.keys(tools).join(', ')}`, isError: true }
  try {
    const args = JSON.parse(call.function.arguments) // models do send broken JSON now and then
    return { output: await run(args), isError: false }
  } catch (err) {
    return { output: `Error: ${err instanceof Error ? err.message : err}`, isError: true }
  }
}

export async function runAgent(prompt: string, maxTurns = 10, maxErrorsInARow = 3): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]
  let errorsInARow = 0

  for (let turn = 1; turn <= maxTurns; turn++) {
    const choice = await callModel(messages)
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const { output, isError } = await runTool(call)
      // Chat Completions has no is_error field: the text itself has to say it failed.
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
      errorsInARow = isError ? errorsInARow + 1 : 0
    }
    // A model stuck on the same wrong move rarely gets out by itself.
    if (errorsInARow >= maxErrorsInARow) throw new Error(`Gave up after ${errorsInARow} tool errors in a row`)
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
// Layer 3 is everything else: a bug in this file, an abort. Nothing catches it here, on purpose.

await runAgent('Open the temple door.')
```

**Python**

```python
# Errors in an agent loop, from scratch. Standard library only, no SDK.
# The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
import json
import random
import time
import urllib.error
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

room = {"lock_oiled": False, "door_open": False}

def use(item, target):
    if item == "oil_can" and target == "lock":
        room["lock_oiled"] = True
        return "You oil the lock. It looks like it might turn now."
    if item == "key" and target == "door":
        # Say what went wrong and what to try instead: this text is all the model will get.
        if not room["lock_oiled"]:
            raise ValueError("the lock is rusted shut and the key won't turn. Oil it first: use oil_can with lock.")
        room["door_open"] = True
        return "Click! The key turns and the door swings open."
    raise ValueError(f"nothing happens when you use {item} with {target}.")

TOOLS = {"look_at": look_at, "use": use}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

def call_model(messages, max_attempts=3):
    """Layer 1, the model's server. Retry only what is likely to pass by itself."""
    for attempt in range(1, max_attempts + 1):
        request = urllib.request.Request(
            f"{LLM['base_url']}/chat/completions",
            data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
            headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
        )
        try:
            with urllib.request.urlopen(request) as response:
                return json.load(response)["choices"][0]
        except urllib.error.HTTPError as err:
            # 429 (too many requests) and 5xx (server trouble) tend to pass.
            # A 400 or a 401 won't: retrying just sends the same broken request again.
            if (err.code != 429 and err.code < 500) or attempt == max_attempts:
                raise
        except urllib.error.URLError:
            # No response at all: the network failed.
            if attempt == max_attempts:
                raise
        # Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
        time.sleep(random.random() * 2 ** (attempt - 1))

def run_tool(call):
    """Layer 2, the tools. Whatever goes wrong becomes a result the model can read."""
    name = call["function"]["name"]
    run = TOOLS.get(name)
    if run is None:
        return f"Unknown tool: {name}. Tools: {', '.join(TOOLS)}", True
    try:
        args = json.loads(call["function"]["arguments"])  # models do send broken JSON now and then
        return run(**args), False
    except Exception as err:
        return f"Error: {err}", True

def run_agent(prompt, max_turns=10, max_errors_in_a_row=3):
    messages = [{"role": "user", "content": prompt}]
    errors_in_a_row = 0

    for _ in range(max_turns):
        choice = call_model(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output, is_error = run_tool(call)
            # Chat Completions has no is_error field: the text itself has to say it failed.
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
            errors_in_a_row = errors_in_a_row + 1 if is_error else 0

        # A model stuck on the same wrong move rarely gets out by itself.
        if errors_in_a_row >= max_errors_in_a_row:
            raise RuntimeError(f"Gave up after {errors_in_a_row} tool errors in a row")

    raise RuntimeError(f"No answer after {max_turns} turns")
# Layer 3 is everything else: a bug in this file, a Ctrl+C. Nothing catches it here, on purpose.

run_agent("Open the temple door.")
```

## 注意事项

- **为模型写错误信息。**说清楚哪里错了、应该怎么做：“锁锈死了。先上油：对 lock 使用 oil_can”。只写“Error 400”，你会再收到一次同样的调用。
- **别泄露模型不该看到的东西。**堆栈跟踪、文件路径和连接字符串都应该进你的日志。模型只需要一句清楚的话。
- **错误也会陷入循环。**模型可能一直重复同一个错误的动作，直到 `maxTurns` 用完。连续出错几次后就停下，或者用钩子拦住它（下一关）。
- **重试要有上限，并加上随机等待。**少量几次尝试，每次翻倍的上限，再加一点随机性（jitter），免得一百个客户端在同一秒一起回来。
- **重试会改变状态的工具时要小心。**如果一次转账或订房的调用超时了，它可能其实已经成功了。再做一次之前，先确认一下。

## 相关模式

- [3 · 设计一个工具](https://harnesspatterns.dev/zh/patterns/designing-a-tool.md)
- [4 · 何时停止](https://harnesspatterns.dev/zh/patterns/when-to-stop.md)
- [6 · 钩子](https://harnesspatterns.dev/zh/patterns/hooks.md)
- [11 · 可观测性与评估](https://harnesspatterns.dev/zh/patterns/observability.md)
