Level 2
Agent Loop
- user
- assistant
- tool_result
EventBus
The problem
A mother bird asks a model "can you clear the pigs' fort?". The model can't fling anything. The best it can do
is reply with a tool request: { name: "launch_red", input: { angle: 40 } }.
If your code makes a single call to the model, the conversation ends right there. You're left with a request nobody ran and no answer. If you run the tool by hand, you hit the same problem on the next round, because the model may need another tool, and then another.
The solution
A loop with a single exit rule:
- Send the model the whole history plus the list of available tools.
- If the reply ends with
stopReason: "end_turn", return it. That's the only normal exit. -
If it ends with
"tool_use", run every requested tool, append the results to the history astool_resultblocks, and go back to step 1.
The model decides what to do; the loop is what does it. That split is the foundation for every other pattern: everything else (steering, context compaction, subagents) hooks into some point of this loop.
The cast
The loop told as a little quest. Once you've met the cast, there's nothing left to decode.
- Astor the loop
- A little tanguero, and the only one who moves. He carries the question to the Oracle, runs to the slingshot for every shot, and brings the answer back to the mother bird.
- The Oracle Provider
- The model. It never touches the slingshot: it only listens to the bandoneón and hands back a note. Orange if it needs a tool, gold if it's done.
- The bandoneón messages[]
- The history, one colored fold per message. The bellows grow every lap, and the Oracle listens to every fold every time. That's the notes floating up to the Oracle.
- The bench ToolRegistry
-
One bird per tool:
launch_red,launch_bomband a third nobody needs today. Astor flings the one the note names and holds up the result: green if it worked. - The mother bird agent.run()
- Your code. It asks the question and waits.
- Trails, score and birds history, tokens, maxTurns
- Every shot leaves its trail in the sky, the way the history keeps every result. The score is tokens, and it jumps more every lap because the whole history is sent again. Every lap costs a bird from the row in the top bar, the turn budget. The numbers are illustrative.
The EventBus panel shows the events the real loop emits at each step of the animation.
The code
With astorlm: The same loop lives in src/agent/loop.ts, with streaming, retries, hooks, parallel tool execution
and cancellation. From the outside, it looks like this.
From scratch: About 40 lines against any OpenAI-compatible endpoint, with no SDK: plain fetch in TypeScript, the
standard library in Python. The three steps above are marked in the comments. Fill in the LLM block
at the top with your own endpoint, model and key.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
// Your game's functions, wrapped as tools: one per bird.
const angle = z.number().min(10).max(80).describe('Launch angle in degrees')
const launchRed = tool({
name: 'launch_red',
description: 'Fling the red bird. Good against wood. Returns what fell and how many pigs are left.',
schema: z.object({ angle }),
execute: async ({ angle }) => level.fling('red', angle), // your code
})
const launchBomb = tool({
name: 'launch_bomb',
description: 'Fling the bomb bird. It explodes on impact: the one to use against stone.',
schema: z.object({ angle }),
execute: async ({ angle }) => level.fling('bomb', angle), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [launchRed, launchBomb],
maxTurns: 10, // the birds in line: a cap on the laps
})
agent.on('tool-start', (tool) => console.log('→', tool.name, tool.input))
const answer = await agent.run('The pigs took our eggs! Can you clear their fort?')
// Agent loop from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { launch_red: launchRed, launch_bomb: launchBomb }
const toolSchemas = [/* one JSON Schema per tool */]
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
// 1. Send the whole history plus the tool list.
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
// 2. No tool calls: the model is done. The only normal exit.
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
// 3. Run each requested tool and feed the result back as a message.
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
if (run) {
try {
output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
# Agent loop from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"launch_red": launch_red, "launch_bomb": launch_bomb}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, max_turns=10):
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
# 1. Send the whole history plus the tool list.
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
# 2. No tool calls: the model is done. The only normal exit.
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
# 3. Run each requested tool and feed the result back as a message.
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
run = TOOLS.get(name)
try:
output = run(**json.loads(call["function"]["arguments"])) if run else f"Unknown tool: {name}"
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
Notice that a tool error doesn't stop the loop: it goes back to the model as text, so it can correct itself on the next round.
When to use it, and what to watch
Whenever the model needs to act (read, search, run something) before it can answer. If all you need is to transform text, a single call is enough, and cheaper.
- Cost grows every lap. Watch the bandoneón and the token counter: the Oracle listens to every fold every turn, so every lap resends everything before it. With enough turns, the context fills up.
-
Always set
maxTurns. A confused model can keep asking for tools forever. Without a limit, the loop runs forever too.