Level 4
When to stop
- user
- assistant
- tool_result
EventBus
The problem
The loop from level 2 has one normal exit: the model replies without asking for a tool. But a model that can't find what it's looking for rarely says so. It tries another search, then another spelling, then another. Each lap resends the whole history, so every lap costs more than the last one.
That's Loopboros, the loop that never ends. maxTurns stops it. But look at world 1-1 again: when
the clock runs out, the loop doesn't fail. It stops and hands back its last message, and that message is a
search request. No answer, no error.
The three exits
-
The flag
end_turnWorld 1-2
The model decides. It replies without asking for any tool. This is the normal way out, and most runs should end here.
Returns: A message with the answer as text.
-
The warp pipe
stopOnToolNamesWorld 1-3
Your code decides, when the model calls a tool you marked. The loop closes right after that tool runs without error. If it fails, the error goes back to the model and the loop continues.
Returns: A message with the tool call. Its input is the answer, already checked against the schema.
-
The clock
maxTurnsWorld 1-1
Nobody decides. The budget runs out. The loop stops after that many turns, whatever the model was doing. It’s a safety net, not an answer.
Returns: Whatever the last message was. Often a tool call.
A terminal tool doesn't save a turn over end_turn. What it gives you is the answer as data, checked
by a schema, ready for your app to use: here, a marker on the level map. Without stopOnToolNames,
after mark_star ran, the loop would ask the model again just so it could say "done".
The code
With astorlm: maxTurns sets the clock and stopOnToolNames marks the warp pipes. run() always
returns the last message, so read it to find out which exit the run took.
From scratch: The loop from level 2 with all three exits marked. Instead of a bare string, it returns how the run ended, so the caller can't mistake a timeout for an answer.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const searchBlocks = tool({
name: 'search_blocks',
description: 'Search the ? blocks in one area of the level. Returns what each one hides.',
schema: z.object({ area: z.string() }),
execute: async ({ area }) => level.search(area), // your code
})
// The answer as data. Its schema is checked before the loop is allowed to stop.
const markStar = tool({
name: 'mark_star',
description: 'Put a marker on the level map where the star is. Call it once, at the end.',
schema: z.object({ item: z.literal('star'), block: z.number().int().positive() }),
execute: async ({ block }) => map.addMarker(block), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [searchBlocks, markStar],
appendSystemPrompt:
'When you know where the star is, deliver it with mark_star. ' +
"If a few searches come back empty, say you couldn't find it.",
maxTurns: 10, // the clock. Without it, the default is 25.
stopOnToolNames: ['mark_star'], // the warp pipe
})
const last = await agent.run('Is there a hidden star in this level?')
// The loop hands back its last message whichever way it ended. Read it to find out which.
const call = last.content.find((block) => block.type === 'tool_use')
if (!call) {
console.log('end_turn:', last.content) // the flag: a plain text answer
} else if (call.name === 'mark_star') {
console.log('answer:', call.input) // the warp pipe: { item, block }, schema-checked
} else {
// The clock: maxTurns ran out while the model was still asking for tools.
throw new Error(`No answer: the run stopped while asking for ${call.name}`)
}
// The three exits of an agent loop, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, unknown>) => Promise<string>
const tools: Record<string, ToolFn> = { search_blocks: searchBlocks, mark_star: markStar }
const toolSchemas = [/* one JSON Schema per tool */]
// Terminal tools: once one runs without error, the loop is over.
const stopOn = new Set(['mark_star'])
// Say how the run ended, so the caller can tell an answer from a timeout.
type Outcome =
| { stop: 'end_turn'; text: string }
| { stop: 'terminal_tool'; input: Record<string, unknown> }
| { stop: 'max_turns'; turns: number }
export async function runAgent(prompt: string, maxTurns = 10): Promise<Outcome> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
// Exit 1, the flag: no tool calls, the model is done.
if (choice.finish_reason !== 'tool_calls') return { stop: 'end_turn', text: reply.content ?? '' }
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let input: Record<string, unknown> = {}
let output = `Unknown tool: ${call.function.name}`
let failed = true
if (run) {
try {
input = JSON.parse(call.function.arguments)
output = await run(input)
failed = false
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
// Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
// If it failed, the error goes back to the model and the loop keeps going.
if (stopOn.has(call.function.name) && !failed) return { stop: 'terminal_tool', input }
}
}
// Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return { stop: 'max_turns', turns: maxTurns }
}
# The three exits of an agent loop, from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"search_blocks": search_blocks, "mark_star": mark_star}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
# Terminal tools: once one runs without error, the loop is over.
STOP_ON = {"mark_star"}
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, max_turns=10):
"""Returns how the run ended, so the caller can tell an answer from a timeout."""
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
# Exit 1, the flag: no tool calls, the model is done.
if choice["finish_reason"] != "tool_calls":
return {"stop": "end_turn", "text": reply.get("content") or ""}
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])
run = TOOLS.get(name)
failed = run is None
try:
output = run(**args) if run else f"Unknown tool: {name}"
except Exception as err:
failed = True
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
# Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
# If it failed, the error goes back to the model and the loop keeps going.
if name in STOP_ON and not failed:
return {"stop": "terminal_tool", "input": args}
# Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return {"stop": "max_turns", "turns": max_turns}
What to watch
-
Always set
maxTurns, and size it to the task. A lookup needs a handful of turns; a refactor may need dozens. The default in astorlm is 25. - Treat the clock as a failure. If the run ended with a tool call, tell the user you couldn't finish, or retry with a clearer prompt. Don't show them an empty answer.
- Give the model a way to give up. Tell it in the system prompt to say "I couldn't find it" when a search keeps coming back empty. A model that is allowed to stop, stops sooner.
-
Turns aren't time.
maxTurnsdoesn't help with a single turn that never ends. For that you need a timeout or anAbortSignal.