Nivel 4
Cuándo parar
- user
- assistant
- tool_result
EventBus
El problema
El bucle del nivel 2 tiene una salida normal: el modelo responde sin pedir una herramienta. Pero un modelo que no encuentra lo que busca rara vez lo dice. Prueba otra búsqueda, después otra forma de escribirlo, después otra más. Cada vuelta reenvía todo el historial, así que cada vuelta cuesta más que la anterior.
Ese es Loopboros, el bucle que nunca termina. maxTurns lo frena. Pero vuelve a mirar el mundo 1-1:
cuando se acaba el reloj, el bucle no falla. Se detiene y te devuelve su último mensaje, y ese mensaje es un
pedido de búsqueda. Ni respuesta, ni error.
Las tres salidas
-
La bandera
end_turnWorld 1-2
Decide el modelo. Responde sin pedir ninguna herramienta. Es la salida normal, y la mayoría de las ejecuciones deberían terminar aquí.
Devuelve: Un mensaje con la respuesta en texto.
-
La tubería de atajo
stopOnToolNamesWorld 1-3
Decide tu código, cuando el modelo llama a una herramienta que marcaste. El bucle se cierra justo después de que esa herramienta se ejecuta sin error. Si falla, el error vuelve al modelo y el bucle sigue.
Devuelve: Un mensaje con la llamada a la herramienta. Su entrada es la respuesta, ya validada contra el esquema.
-
El reloj
maxTurnsWorld 1-1
No decide nadie. Se acaba el presupuesto. El bucle se detiene después de esa cantidad de turnos, haga lo que haga el modelo. Es una red de seguridad, no una respuesta.
Devuelve: Lo que haya sido el último mensaje. A menudo, una llamada a herramienta.
Una herramienta terminal no te ahorra un turno frente a end_turn. Lo que te da es la respuesta como
datos, validados por un esquema, listos para que los use tu app: aquí, un marcador en el mapa del nivel. Sin
stopOnToolNames, después de ejecutar mark_star el bucle volvería a consultar al modelo
solo para que dijera "listo".
El código
Con astorlm: maxTurns pone el reloj y stopOnToolNames marca las
tuberías de atajo. run() siempre devuelve el último mensaje, así que léelo para saber por qué
salida terminó la ejecución.
Desde cero: El bucle del nivel 2 con las tres salidas marcadas. En lugar de un string pelado, devuelve cómo terminó la ejecución, para que quien lo llama no confunda un timeout con una respuesta.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const searchBlocks = tool({
name: 'search_blocks',
description: 'Search the ? blocks in one area of the level. Returns what each one hides.',
schema: z.object({ area: z.string() }),
execute: async ({ area }) => level.search(area), // your code
})
// The answer as data. Its schema is checked before the loop is allowed to stop.
const markStar = tool({
name: 'mark_star',
description: 'Put a marker on the level map where the star is. Call it once, at the end.',
schema: z.object({ item: z.literal('star'), block: z.number().int().positive() }),
execute: async ({ block }) => map.addMarker(block), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [searchBlocks, markStar],
appendSystemPrompt:
'When you know where the star is, deliver it with mark_star. ' +
"If a few searches come back empty, say you couldn't find it.",
maxTurns: 10, // the clock. Without it, the default is 25.
stopOnToolNames: ['mark_star'], // the warp pipe
})
const last = await agent.run('Is there a hidden star in this level?')
// The loop hands back its last message whichever way it ended. Read it to find out which.
const call = last.content.find((block) => block.type === 'tool_use')
if (!call) {
console.log('end_turn:', last.content) // the flag: a plain text answer
} else if (call.name === 'mark_star') {
console.log('answer:', call.input) // the warp pipe: { item, block }, schema-checked
} else {
// The clock: maxTurns ran out while the model was still asking for tools.
throw new Error(`No answer: the run stopped while asking for ${call.name}`)
}
// The three exits of an agent loop, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, unknown>) => Promise<string>
const tools: Record<string, ToolFn> = { search_blocks: searchBlocks, mark_star: markStar }
const toolSchemas = [/* one JSON Schema per tool */]
// Terminal tools: once one runs without error, the loop is over.
const stopOn = new Set(['mark_star'])
// Say how the run ended, so the caller can tell an answer from a timeout.
type Outcome =
| { stop: 'end_turn'; text: string }
| { stop: 'terminal_tool'; input: Record<string, unknown> }
| { stop: 'max_turns'; turns: number }
export async function runAgent(prompt: string, maxTurns = 10): Promise<Outcome> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
// Exit 1, the flag: no tool calls, the model is done.
if (choice.finish_reason !== 'tool_calls') return { stop: 'end_turn', text: reply.content ?? '' }
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let input: Record<string, unknown> = {}
let output = `Unknown tool: ${call.function.name}`
let failed = true
if (run) {
try {
input = JSON.parse(call.function.arguments)
output = await run(input)
failed = false
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
// Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
// If it failed, the error goes back to the model and the loop keeps going.
if (stopOn.has(call.function.name) && !failed) return { stop: 'terminal_tool', input }
}
}
// Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return { stop: 'max_turns', turns: maxTurns }
}
# The three exits of an agent loop, from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"search_blocks": search_blocks, "mark_star": mark_star}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
# Terminal tools: once one runs without error, the loop is over.
STOP_ON = {"mark_star"}
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, max_turns=10):
"""Returns how the run ended, so the caller can tell an answer from a timeout."""
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
# Exit 1, the flag: no tool calls, the model is done.
if choice["finish_reason"] != "tool_calls":
return {"stop": "end_turn", "text": reply.get("content") or ""}
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])
run = TOOLS.get(name)
failed = run is None
try:
output = run(**args) if run else f"Unknown tool: {name}"
except Exception as err:
failed = True
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
# Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
# If it failed, the error goes back to the model and the loop keeps going.
if name in STOP_ON and not failed:
return {"stop": "terminal_tool", "input": args}
# Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return {"stop": "max_turns", "turns": max_turns}
Qué vigilar
-
Define siempre
maxTurns, y dimensiónalo según la tarea. Una consulta necesita un puñado de turnos; un refactor puede necesitar docenas. El valor por defecto en astorlm es 25. - Trata el reloj como una falla. Si la ejecución terminó con una llamada a herramienta, dile al usuario que no pudiste terminar, o reintenta con un prompt más claro. No le muestres una respuesta vacía.
- Dale al modelo una forma de rendirse. Dile en el system prompt que diga "no lo encontré" cuando una búsqueda vuelve vacía una y otra vez. Un modelo al que se le permite parar, para antes.
-
Los turnos no son tiempo.
maxTurnsno ayuda con un solo turno que nunca termina. Para eso necesitas un timeout o unAbortSignal.