> Level 6 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/hooks · All patterns: https://harnesspatterns.dev/llms.txt

# Hooks

Events let you watch the loop. Hooks let you change it. A hook is a function of yours that the loop calls at a fixed point, and whatever it returns, the loop obeys.

## The problem

Your travel assistant works. Then the company adds a rule: no first-class tickets without a manager’s approval. And legal adds another: the passenger’s ID number must never reach the model.

Neither rule is about the model. The model can be told, but a prompt is a request, not a lock. Both rules are about what the *loop* does: which tool calls it runs, and what it puts into the history. If the loop gives you no way in, the only option left is to copy its code and edit it. That’s Ironclad, the sealed loop.

Events don’t help either. An event tells you that `book_ticket` is about to run. By the time your listener gets it, nothing you do there can stop it.

## The solution

The loop calls your functions at fixed points of every turn, and uses what they return. Five points cover almost everything:

- `beforeTurn`
   **At the start of every turn.**
   Check a budget, log the turn, stop a run that has gone on too long.
- `beforeProviderCall`
   **Right before the request goes to the model.**
   Change what gets sent: trim old messages, add today’s date, hide a tool this turn.
- `beforeToolExecution`
   **After the model asks for a tool, before it runs.**
   Let it through, refuse it (the model gets your reason instead), or answer with a canned result.
- `afterToolExecution`
   **After the tool runs, before the result joins the history.**
   Rewrite what the model will read: hide personal data, shorten a huge output.
- `afterTurn`
   **Once the model’s reply and any tool results are in.**
   Save progress, update a dashboard, count the cost.

The rule of thumb: **events watch, hooks change.** Use an event when you only want to know what happened. Use a hook when you need to decide what happens.

## The cast

Same cast as always, on a model railway this time.

- **The circuit** (the loop): A closed ring of track that only runs one way. Every lap is a turn: past the Oracle, past the tools, round again.
- **Astor** (the loop's runner): Pumps the handcar around the circuit, with the bandoneón of messages on his back.
- **The booths** (hooks): One per hook point. An empty booth does nothing. A staffed one stops the handcar, checks what it carries, and can lower its barrier or stamp over the cargo. This run staffs two: `beforeToolExecution` and `afterToolExecution`.
- **The stands** (EventBus): Three spectators who write down everything that passes. They see it all, and can touch none of it.
- **The platforms** (tools): `find_trains` and `book_ticket`, on the far curve.

In the EventBus panel, the `hook` and `code` lines mark your own functions running: your hooks and your tools. astorlm doesn’t emit events for them. Notice where they fall: `tool_execution_end` comes after `afterToolExecution`, so it already carries the stamped text.

## The code

**With astorlm:** Pass a `hooks` object to the agent. `beforeToolExecution` returns `{ authorize: false }` to refuse a call, and `afterToolExecution` returns the text the model will read.

**From scratch:** The loop from level 2, with a call to each of the five hooks. A hook nobody set is just skipped.

**With astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const findTrains = tool({
  name: 'find_trains',
  description: 'List the trains to a destination on a date, with the fare for each class.',
  schema: z.object({ to: z.string(), date: z.string() }),
  execute: async ({ to, date }) => searchTimetable(to, date), // your code
})

const bookTicket = tool({
  name: 'book_ticket',
  description: 'Book one seat on a train for the employee who is asking.',
  schema: z.object({ train: z.number().int(), seat_class: z.enum(['first', 'tourist']) }),
  execute: async ({ train, seat_class }) => reserveSeat(train, seat_class), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [findTrains, bookTicket],
  maxTurns: 10,
  hooks: {
    // The first booth: runs before every tool call, and decides whether it runs at all.
    beforeToolExecution: async ({ toolName, input }) => {
      const { seat_class } = input as { seat_class?: string }
      if (toolName === 'book_ticket' && seat_class === 'first') {
        // The tool never runs. The model reads this text as an error result instead.
        return { authorize: false, mockResult: 'Blocked by policy: first class needs a manager’s approval. Book tourist instead.' }
      }
      return { authorize: true }
    },
    // The second booth: runs after every tool call. What you return is what the model reads.
    afterToolExecution: async ({ output }) => output.replace(/DNI [\d.]+/g, 'DNI ***'),
  },
})

// Events only watch. By the time this fires, the hook has already stamped over the DNI.
agent.on('tool-end', ({ name, output, isError }) => console.log(name, isError ? 'refused:' : 'ok:', output))

const last = await agent.run('Book me the most comfortable seat to Mar del Plata on Friday.')
console.log(last.content)
```

**TypeScript**

```ts
// The agent loop with hooks, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

type Args = Record<string, string | number>
type ToolFn = (args: Args) => Promise<string>
const tools: Record<string, ToolFn> = { find_trains: findTrains, book_ticket: bookTicket }
const toolSchemas = [/* one JSON Schema per tool */]

// The five points where the loop lets your code in. Every one is optional.
type Hooks = {
  beforeTurn?: (turn: number, messages: Message[]) => Promise<void>
  // Return the messages to send: trim them, add context, or pass them through.
  beforeProviderCall?: (messages: Message[]) => Promise<Message[]>
  // Say no, and the tool never runs: `result` goes back to the model instead.
  beforeToolExecution?: (name: string, args: Args) => Promise<{ authorize: boolean; result?: string }>
  // Whatever you return is what the model reads.
  afterToolExecution?: (name: string, output: string) => Promise<string>
  afterTurn?: (turn: number, reply: Message) => Promise<void>
}

export async function runAgent(prompt: string, hooks: Hooks = {}, maxTurns = 10): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]

  for (let turn = 1; turn <= maxTurns; turn++) {
    await hooks.beforeTurn?.(turn, messages)
    const outgoing = (await hooks.beforeProviderCall?.(messages)) ?? messages

    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages: outgoing, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)

    if (choice.finish_reason !== 'tool_calls') {
      await hooks.afterTurn?.(turn, reply)
      return reply.content ?? ''
    }

    for (const call of reply.tool_calls ?? []) {
      const name = call.function.name
      const run = tools[name]
      let output = `Unknown tool: ${name}`
      try {
        const args: Args = JSON.parse(call.function.arguments)
        // Booth 1: before the tool runs.
        const gate = (await hooks.beforeToolExecution?.(name, args)) ?? { authorize: true }
        if (!gate.authorize) output = gate.result ?? 'Rejected by policy.'
        else if (run) output = await run(args)
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      // Booth 2: before the result joins the history.
      output = (await hooks.afterToolExecution?.(name, output)) ?? output
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
    await hooks.afterTurn?.(turn, reply)
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// The two booths from the animation.
const answer = await runAgent('Book me the most comfortable seat to Mar del Plata on Friday.', {
  beforeToolExecution: async (name, args) =>
    name === 'book_ticket' && args.seat_class === 'first'
      ? { authorize: false, result: 'Blocked by policy: first class needs a manager’s approval. Book tourist instead.' }
      : { authorize: true },
  afterToolExecution: async (_name, output) => output.replace(/DNI [\d.]+/g, 'DNI ***'),
})
```

**Python**

```python
# The agent loop with hooks, from scratch. Standard library only, no SDK.
import json
import re
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"find_trains": find_trains, "book_ticket": book_ticket}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

# The five points where the loop lets your code in. Every one is optional:
#   before_turn(turn, messages)
#   before_provider_call(messages) -> the messages to send
#   before_tool_execution(name, args) -> {"authorize": bool, "result": str}
#   after_tool_execution(name, output) -> the text the model will read
#   after_turn(turn, reply)

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

def run_agent(prompt, hooks=None, max_turns=10):
    hooks = hooks or {}

    def call_hook(point, *args):
        return hooks[point](*args) if point in hooks else None

    messages = [{"role": "user", "content": prompt}]

    for turn in range(1, max_turns + 1):
        call_hook("before_turn", turn, messages)
        outgoing = call_hook("before_provider_call", messages) or messages

        choice = chat(outgoing)
        reply = choice["message"]
        messages.append(reply)

        if choice["finish_reason"] != "tool_calls":
            call_hook("after_turn", turn, reply)
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            run = TOOLS.get(name)
            try:
                args = json.loads(call["function"]["arguments"])
                # Booth 1: before the tool runs. Say no, and it never does.
                gate = call_hook("before_tool_execution", name, args) or {"authorize": True}
                if not gate["authorize"]:
                    output = gate.get("result", "Rejected by policy.")
                else:
                    output = run(**args) if run else f"Unknown tool: {name}"
            except Exception as err:
                output = f"Error: {err}"
            # Booth 2: before the result joins the history. What it returns is what the model reads.
            output = call_hook("after_tool_execution", name, output) or output
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

        call_hook("after_turn", turn, reply)

    raise RuntimeError(f"No answer after {max_turns} turns")

# The two booths from the animation.
def check_policy(name, args):
    if name == "book_ticket" and args.get("seat_class") == "first":
        return {"authorize": False, "result": "Blocked by policy: first class needs a manager's approval. Book tourist instead."}
    return {"authorize": True}

def hide_ids(name, output):
    return re.sub(r"DNI [\d.]+", "DNI ***", output)

answer = run_agent(
    "Book me the most comfortable seat to Mar del Plata on Friday.",
    hooks={"before_tool_execution": check_policy, "after_tool_execution": hide_ids},
)
```

## What to watch

- **Say why when you refuse.** The refusal goes back to the model as an error result. "Blocked by policy: book tourist instead" gets you a tourist ticket. A bare "denied" gets you the same call again.
- **Hooks run on every call, so keep them fast.** A hook that queries a database adds that delay to every tool and every turn.
- **A hook that throws takes the run down.** The loop catches errors from your tools, not from your hooks. Wrap anything that can fail.
- **Don’t use a hook to watch.** If you only log, listen to events. Keep hooks for the times you need to change something.

## Related patterns

- [2 · Agent Loop](https://harnesspatterns.dev/patterns/agent-loop.md)
- [5 · Errors in the loop](https://harnesspatterns.dev/patterns/errors-in-the-loop.md)
- [7 · The backpack fills up](https://harnesspatterns.dev/patterns/compaction.md)
- [13 · Human in the loop](https://harnesspatterns.dev/patterns/human-in-the-loop.md)
- [14 · Security and sandboxing](https://harnesspatterns.dev/patterns/security.md)
