Level 6
Hooks
- user
- assistant
- tool_result
- tool_result (error)
EventBus
The problem
Your travel assistant works. Then the company adds a rule: no first-class tickets without a manager’s approval. And legal adds another: the passenger’s ID number must never reach the model.
Neither rule is about the model. The model can be told, but a prompt is a request, not a lock. Both rules are about what the loop does: which tool calls it runs, and what it puts into the history. If the loop gives you no way in, the only option left is to copy its code and edit it. That’s Ironclad, the sealed loop.
Events don’t help either. An event tells you that book_ticket is about to run. By the time your
listener gets it, nothing you do there can stop it.
The solution
The loop calls your functions at fixed points of every turn, and uses what they return. Five points cover almost everything:
-
beforeTurnAt the start of every turn.
Check a budget, log the turn, stop a run that has gone on too long.
-
beforeProviderCallRight before the request goes to the model.
Change what gets sent: trim old messages, add today’s date, hide a tool this turn.
-
beforeToolExecutionAfter the model asks for a tool, before it runs.
Let it through, refuse it (the model gets your reason instead), or answer with a canned result.
-
afterToolExecutionAfter the tool runs, before the result joins the history.
Rewrite what the model will read: hide personal data, shorten a huge output.
-
afterTurnOnce the model’s reply and any tool results are in.
Save progress, update a dashboard, count the cost.
The rule of thumb: events watch, hooks change. Use an event when you only want to know what happened. Use a hook when you need to decide what happens.
The cast
Same cast as always, on a model railway this time.
- The circuit the loop
- A closed ring of track that only runs one way. Every lap is a turn: past the Oracle, past the tools, round again.
- Astor the loop's runner
- Pumps the handcar around the circuit, with the bandoneón of messages on his back.
- The booths hooks
-
One per hook point. An empty booth does nothing. A staffed one stops the handcar, checks what it carries, and can
lower its barrier or stamp over the cargo. This run staffs two:
beforeToolExecutionandafterToolExecution. - The stands EventBus
- Three spectators who write down everything that passes. They see it all, and can touch none of it.
- The platforms tools
find_trainsandbook_ticket, on the far curve.
In the EventBus panel, the hook and code lines mark your own functions running: your
hooks and your tools. astorlm doesn’t emit events for them. Notice where they fall:
tool_execution_end comes after afterToolExecution, so it already carries the stamped
text.
The code
With astorlm: Pass a hooks object to the agent. beforeToolExecution
returns { authorize: false } to refuse a call, and afterToolExecution returns the
text the model will read.
From scratch: The loop from level 2, with a call to each of the five hooks. A hook nobody set is just skipped.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const findTrains = tool({
name: 'find_trains',
description: 'List the trains to a destination on a date, with the fare for each class.',
schema: z.object({ to: z.string(), date: z.string() }),
execute: async ({ to, date }) => searchTimetable(to, date), // your code
})
const bookTicket = tool({
name: 'book_ticket',
description: 'Book one seat on a train for the employee who is asking.',
schema: z.object({ train: z.number().int(), seat_class: z.enum(['first', 'tourist']) }),
execute: async ({ train, seat_class }) => reserveSeat(train, seat_class), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [findTrains, bookTicket],
maxTurns: 10,
hooks: {
// The first booth: runs before every tool call, and decides whether it runs at all.
beforeToolExecution: async ({ toolName, input }) => {
const { seat_class } = input as { seat_class?: string }
if (toolName === 'book_ticket' && seat_class === 'first') {
// The tool never runs. The model reads this text as an error result instead.
return { authorize: false, mockResult: 'Blocked by policy: first class needs a manager’s approval. Book tourist instead.' }
}
return { authorize: true }
},
// The second booth: runs after every tool call. What you return is what the model reads.
afterToolExecution: async ({ output }) => output.replace(/DNI [\d.]+/g, 'DNI ***'),
},
})
// Events only watch. By the time this fires, the hook has already stamped over the DNI.
agent.on('tool-end', ({ name, output, isError }) => console.log(name, isError ? 'refused:' : 'ok:', output))
const last = await agent.run('Book me the most comfortable seat to Mar del Plata on Friday.')
console.log(last.content)
// The agent loop with hooks, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type Args = Record<string, string | number>
type ToolFn = (args: Args) => Promise<string>
const tools: Record<string, ToolFn> = { find_trains: findTrains, book_ticket: bookTicket }
const toolSchemas = [/* one JSON Schema per tool */]
// The five points where the loop lets your code in. Every one is optional.
type Hooks = {
beforeTurn?: (turn: number, messages: Message[]) => Promise<void>
// Return the messages to send: trim them, add context, or pass them through.
beforeProviderCall?: (messages: Message[]) => Promise<Message[]>
// Say no, and the tool never runs: `result` goes back to the model instead.
beforeToolExecution?: (name: string, args: Args) => Promise<{ authorize: boolean; result?: string }>
// Whatever you return is what the model reads.
afterToolExecution?: (name: string, output: string) => Promise<string>
afterTurn?: (turn: number, reply: Message) => Promise<void>
}
export async function runAgent(prompt: string, hooks: Hooks = {}, maxTurns = 10): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
await hooks.beforeTurn?.(turn, messages)
const outgoing = (await hooks.beforeProviderCall?.(messages)) ?? messages
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages: outgoing, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') {
await hooks.afterTurn?.(turn, reply)
return reply.content ?? ''
}
for (const call of reply.tool_calls ?? []) {
const name = call.function.name
const run = tools[name]
let output = `Unknown tool: ${name}`
try {
const args: Args = JSON.parse(call.function.arguments)
// Booth 1: before the tool runs.
const gate = (await hooks.beforeToolExecution?.(name, args)) ?? { authorize: true }
if (!gate.authorize) output = gate.result ?? 'Rejected by policy.'
else if (run) output = await run(args)
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
// Booth 2: before the result joins the history.
output = (await hooks.afterToolExecution?.(name, output)) ?? output
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
await hooks.afterTurn?.(turn, reply)
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// The two booths from the animation.
const answer = await runAgent('Book me the most comfortable seat to Mar del Plata on Friday.', {
beforeToolExecution: async (name, args) =>
name === 'book_ticket' && args.seat_class === 'first'
? { authorize: false, result: 'Blocked by policy: first class needs a manager’s approval. Book tourist instead.' }
: { authorize: true },
afterToolExecution: async (_name, output) => output.replace(/DNI [\d.]+/g, 'DNI ***'),
})
# The agent loop with hooks, from scratch. Standard library only, no SDK.
import json
import re
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"find_trains": find_trains, "book_ticket": book_ticket}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
# The five points where the loop lets your code in. Every one is optional:
# before_turn(turn, messages)
# before_provider_call(messages) -> the messages to send
# before_tool_execution(name, args) -> {"authorize": bool, "result": str}
# after_tool_execution(name, output) -> the text the model will read
# after_turn(turn, reply)
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, hooks=None, max_turns=10):
hooks = hooks or {}
def call_hook(point, *args):
return hooks[point](*args) if point in hooks else None
messages = [{"role": "user", "content": prompt}]
for turn in range(1, max_turns + 1):
call_hook("before_turn", turn, messages)
outgoing = call_hook("before_provider_call", messages) or messages
choice = chat(outgoing)
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
call_hook("after_turn", turn, reply)
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
run = TOOLS.get(name)
try:
args = json.loads(call["function"]["arguments"])
# Booth 1: before the tool runs. Say no, and it never does.
gate = call_hook("before_tool_execution", name, args) or {"authorize": True}
if not gate["authorize"]:
output = gate.get("result", "Rejected by policy.")
else:
output = run(**args) if run else f"Unknown tool: {name}"
except Exception as err:
output = f"Error: {err}"
# Booth 2: before the result joins the history. What it returns is what the model reads.
output = call_hook("after_tool_execution", name, output) or output
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
call_hook("after_turn", turn, reply)
raise RuntimeError(f"No answer after {max_turns} turns")
# The two booths from the animation.
def check_policy(name, args):
if name == "book_ticket" and args.get("seat_class") == "first":
return {"authorize": False, "result": "Blocked by policy: first class needs a manager's approval. Book tourist instead."}
return {"authorize": True}
def hide_ids(name, output):
return re.sub(r"DNI [\d.]+", "DNI ***", output)
answer = run_agent(
"Book me the most comfortable seat to Mar del Plata on Friday.",
hooks={"before_tool_execution": check_policy, "after_tool_execution": hide_ids},
)
What to watch
- Say why when you refuse. The refusal goes back to the model as an error result. "Blocked by policy: book tourist instead" gets you a tourist ticket. A bare "denied" gets you the same call again.
- Hooks run on every call, so keep them fast. A hook that queries a database adds that delay to every tool and every turn.
- A hook that throws takes the run down. The loop catches errors from your tools, not from your hooks. Wrap anything that can fail.
- Don’t use a hook to watch. If you only log, listen to events. Keep hooks for the times you need to change something.