> Agent Harness Patterns 第 6 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/hooks · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 钩子

事件让你观察循环。钩子让你改变循环。钩子就是你写的一个函数，循环会在固定的点调用它，而它返回什么，循环就照做什么。

## 问题

你的差旅助手运转正常。然后公司加了一条规定：没有经理批准，不许订头等舱。法务又加了一条：乘客的身份证号码绝不能到达模型那里。

这两条规定都和模型无关。你可以告诉模型，但提示词只是一个请求，不是一把锁。这两条规定管的是*循环*做什么：它执行哪些工具调用，往历史记录里放什么。如果循环不给你任何介入的入口，剩下的唯一办法就是复制它的代码再改。这就是 Ironclad，那个密封的循环。

事件也帮不上忙。事件只能告诉你 `book_ticket` 马上就要运行了。等你的监听器收到它时，你在那里做什么都拦不住了。

## 解决方案

循环在每一轮的固定点调用你的函数，并使用它们的返回值。五个点几乎涵盖了一切：

- `beforeTurn`
   **每一轮开始时。**
   检查预算、记录这一轮、停掉一次跑得太久的运行。
- `beforeProviderCall`
   **请求发往模型之前的那一刻。**
   修改要发送的内容：裁掉旧消息、加上今天的日期、在这一轮隐藏某个工具。
- `beforeToolExecution`
   **模型申请工具之后、工具运行之前。**
   放行、拒绝（模型会收到你给出的理由），或者用一个预设结果来回应。
- `afterToolExecution`
   **工具运行之后、结果进入历史记录之前。**
   改写模型将要读到的内容：隐藏个人数据、截短超长的输出。
- `afterTurn`
   **模型的回复和所有工具结果都到齐之后。**
   保存进度、更新仪表盘、统计成本。

经验法则：**事件负责观察，钩子负责改变。**只想知道发生了什么时，用事件。需要决定会发生什么时，用钩子。

## 角色

还是那群熟悉的角色，这次换到了铁路模型上。

- **环形轨道** (循环): 一圈只朝一个方向跑的封闭轨道。每跑一圈就是一轮：经过神谕者，经过工具，再来一圈。
- **Astor** (跑循环的人): 背着装满消息的手风琴，压着手摇车绕轨道跑。
- **岗亭** (钩子): 每个钩子点一座。没人值守的岗亭什么也不做。有人值守的岗亭会拦下手摇车，检查它载着什么，可以放下栏杆，也可以在货物上盖章遮挡。这次运行有两座岗亭有人值守：`beforeToolExecution` 和 `afterToolExecution`。
- **看台** (EventBus): 三位观众，把经过的一切都记下来。他们什么都看得到，却什么都碰不了。
- **站台** (工具): 远处弯道上的 `find_trains` 和 `book_ticket`。

在 EventBus 面板里，`hook` 和 `code` 两种行标记的是你自己的函数在运行：你的钩子和你的工具。astorlm 不会为它们发出事件。注意它们出现的位置：`tool_execution_end` 出现在 `afterToolExecution` 之后，所以它带的已经是盖过章的文本。

## 代码

**使用 astorlm：**给智能体传一个 `hooks` 对象。`beforeToolExecution` 返回 `{ authorize: false }` 来拒绝一次调用，`afterToolExecution` 返回模型将要读到的文本。

**从零手写：**第 2 关的循环，加上对五个钩子各自的一次调用。没人设置的钩子会被直接跳过。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const findTrains = tool({
  name: 'find_trains',
  description: 'List the trains to a destination on a date, with the fare for each class.',
  schema: z.object({ to: z.string(), date: z.string() }),
  execute: async ({ to, date }) => searchTimetable(to, date), // your code
})

const bookTicket = tool({
  name: 'book_ticket',
  description: 'Book one seat on a train for the employee who is asking.',
  schema: z.object({ train: z.number().int(), seat_class: z.enum(['first', 'tourist']) }),
  execute: async ({ train, seat_class }) => reserveSeat(train, seat_class), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [findTrains, bookTicket],
  maxTurns: 10,
  hooks: {
    // The first booth: runs before every tool call, and decides whether it runs at all.
    beforeToolExecution: async ({ toolName, input }) => {
      const { seat_class } = input as { seat_class?: string }
      if (toolName === 'book_ticket' && seat_class === 'first') {
        // The tool never runs. The model reads this text as an error result instead.
        return { authorize: false, mockResult: 'Blocked by policy: first class needs a manager’s approval. Book tourist instead.' }
      }
      return { authorize: true }
    },
    // The second booth: runs after every tool call. What you return is what the model reads.
    afterToolExecution: async ({ output }) => output.replace(/DNI [\d.]+/g, 'DNI ***'),
  },
})

// Events only watch. By the time this fires, the hook has already stamped over the DNI.
agent.on('tool-end', ({ name, output, isError }) => console.log(name, isError ? 'refused:' : 'ok:', output))

const last = await agent.run('Book me the most comfortable seat to Mar del Plata on Friday.')
console.log(last.content)
```

**TypeScript**

```ts
// The agent loop with hooks, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

type Args = Record<string, string | number>
type ToolFn = (args: Args) => Promise<string>
const tools: Record<string, ToolFn> = { find_trains: findTrains, book_ticket: bookTicket }
const toolSchemas = [/* one JSON Schema per tool */]

// The five points where the loop lets your code in. Every one is optional.
type Hooks = {
  beforeTurn?: (turn: number, messages: Message[]) => Promise<void>
  // Return the messages to send: trim them, add context, or pass them through.
  beforeProviderCall?: (messages: Message[]) => Promise<Message[]>
  // Say no, and the tool never runs: `result` goes back to the model instead.
  beforeToolExecution?: (name: string, args: Args) => Promise<{ authorize: boolean; result?: string }>
  // Whatever you return is what the model reads.
  afterToolExecution?: (name: string, output: string) => Promise<string>
  afterTurn?: (turn: number, reply: Message) => Promise<void>
}

export async function runAgent(prompt: string, hooks: Hooks = {}, maxTurns = 10): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]

  for (let turn = 1; turn <= maxTurns; turn++) {
    await hooks.beforeTurn?.(turn, messages)
    const outgoing = (await hooks.beforeProviderCall?.(messages)) ?? messages

    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages: outgoing, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)

    if (choice.finish_reason !== 'tool_calls') {
      await hooks.afterTurn?.(turn, reply)
      return reply.content ?? ''
    }

    for (const call of reply.tool_calls ?? []) {
      const name = call.function.name
      const run = tools[name]
      let output = `Unknown tool: ${name}`
      try {
        const args: Args = JSON.parse(call.function.arguments)
        // Booth 1: before the tool runs.
        const gate = (await hooks.beforeToolExecution?.(name, args)) ?? { authorize: true }
        if (!gate.authorize) output = gate.result ?? 'Rejected by policy.'
        else if (run) output = await run(args)
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      // Booth 2: before the result joins the history.
      output = (await hooks.afterToolExecution?.(name, output)) ?? output
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
    await hooks.afterTurn?.(turn, reply)
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// The two booths from the animation.
const answer = await runAgent('Book me the most comfortable seat to Mar del Plata on Friday.', {
  beforeToolExecution: async (name, args) =>
    name === 'book_ticket' && args.seat_class === 'first'
      ? { authorize: false, result: 'Blocked by policy: first class needs a manager’s approval. Book tourist instead.' }
      : { authorize: true },
  afterToolExecution: async (_name, output) => output.replace(/DNI [\d.]+/g, 'DNI ***'),
})
```

**Python**

```python
# The agent loop with hooks, from scratch. Standard library only, no SDK.
import json
import re
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"find_trains": find_trains, "book_ticket": book_ticket}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

# The five points where the loop lets your code in. Every one is optional:
#   before_turn(turn, messages)
#   before_provider_call(messages) -> the messages to send
#   before_tool_execution(name, args) -> {"authorize": bool, "result": str}
#   after_tool_execution(name, output) -> the text the model will read
#   after_turn(turn, reply)

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

def run_agent(prompt, hooks=None, max_turns=10):
    hooks = hooks or {}

    def call_hook(point, *args):
        return hooks[point](*args) if point in hooks else None

    messages = [{"role": "user", "content": prompt}]

    for turn in range(1, max_turns + 1):
        call_hook("before_turn", turn, messages)
        outgoing = call_hook("before_provider_call", messages) or messages

        choice = chat(outgoing)
        reply = choice["message"]
        messages.append(reply)

        if choice["finish_reason"] != "tool_calls":
            call_hook("after_turn", turn, reply)
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            run = TOOLS.get(name)
            try:
                args = json.loads(call["function"]["arguments"])
                # Booth 1: before the tool runs. Say no, and it never does.
                gate = call_hook("before_tool_execution", name, args) or {"authorize": True}
                if not gate["authorize"]:
                    output = gate.get("result", "Rejected by policy.")
                else:
                    output = run(**args) if run else f"Unknown tool: {name}"
            except Exception as err:
                output = f"Error: {err}"
            # Booth 2: before the result joins the history. What it returns is what the model reads.
            output = call_hook("after_tool_execution", name, output) or output
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

        call_hook("after_turn", turn, reply)

    raise RuntimeError(f"No answer after {max_turns} turns")

# The two booths from the animation.
def check_policy(name, args):
    if name == "book_ticket" and args.get("seat_class") == "first":
        return {"authorize": False, "result": "Blocked by policy: first class needs a manager's approval. Book tourist instead."}
    return {"authorize": True}

def hide_ids(name, output):
    return re.sub(r"DNI [\d.]+", "DNI ***", output)

answer = run_agent(
    "Book me the most comfortable seat to Mar del Plata on Friday.",
    hooks={"before_tool_execution": check_policy, "after_tool_execution": hide_ids},
)
```

## 注意事项

- **拒绝时要说明原因。**拒绝会以错误结果的形式回到模型那里。“因政策被拦截：请改订经济舱”能换来一张经济舱车票。一句光秃秃的“拒绝”只会换来同样的调用再来一次。
- **钩子在每次调用时都会运行，所以要让它们足够快。**一个要查数据库的钩子，会把这段延迟加到每个工具、每一轮上。
- **抛出异常的钩子会让整次运行崩掉。**循环会接住你工具里的错误，但不会接住你钩子里的错误。任何可能出错的地方都要包起来。
- **别用钩子来观察。**如果你只是记日志，就去监听事件。把钩子留给你需要改变某些东西的时候。

## 相关模式

- [2 · 智能体循环](https://harnesspatterns.dev/zh/patterns/agent-loop.md)
- [5 · 循环中的错误](https://harnesspatterns.dev/zh/patterns/errors-in-the-loop.md)
- [7 · 背包装满了](https://harnesspatterns.dev/zh/patterns/compaction.md)
- [13 · 人在回路](https://harnesspatterns.dev/zh/patterns/human-in-the-loop.md)
- [14 · 安全与沙箱](https://harnesspatterns.dev/zh/patterns/security.md)
