Nível 5
Erros no loop
- user
- assistant
- tool_result
- tool_result (erro)
EventBus
O problema
Um agente joga um jogo de aventura. As ferramentas dele são os verbos na parte de baixo da tela,
look_at e use, e a tarefa é "abra a porta do templo". O modelo faz o óbvio: usar a
chave na porta. Mas a fechadura está enferrujada, e a ferramenta lança uma exceção.
Os jogos de aventura antigos se dividiam em duas escolas diante de momentos assim. Em alguns, um movimento errado te matava: game over, de volta ao último save. Em outros, você não podia morrer; o jogo só dizia por que aquilo não funcionava, e você tentava outra coisa. Um loop de agente também precisa escolher uma escola.
Se o loop não captura o erro, o Krash vence: a exceção sobe pelo loop e o run() a lança. A execução
se perde, e o erro dizia exatamente o que fazer: "lubrifique primeiro". O único leitor que podia usar essa
mensagem nunca a recebeu. Capture-o, devolva-o como resultado de ferramenta, e o modelo lê e tenta de novo. Isso
custa um turno. Não capturar custa a execução inteira.
Três tipos de falha
-
O modelo pediu errado
Entrada inválida, JSON quebrado, uma ferramenta que não existe, uma chave que não gira.
O modelo. O erro volta como um resultado de ferramenta marcado como erro. O modelo lê e tenta outra coisa no turno seguinte.
-
O servidor do modelo falhou
429 (requisições demais), 5xx (problema no servidor), uma conexão que caiu.
O seu código, tentando de novo. Espere um pouco e chame de novo, um pouco mais a cada vez. O modelo nunca fica sabendo: não houve turno para ler.
-
Ninguém consegue resolver
Um 401 (API key errada), um bug no seu código, o usuário cancelou.
Ninguém. Deixe a exceção subir. Tentar de novo não vai ajudar, e esconder isso do modelo só o faz chutar.
O segredo é não misturá-las. Tentar de novo um 400 manda a mesma requisição quebrada outra vez. Mostrar um 503 ao modelo gasta um turno com algo que ele não pode resolver. E capturar um bug do seu próprio código só o esconde.
O código
Com astorlm: As ferramentas simplesmente lançam exceções, e o astorlm faz a captura. Você liga o
retry para o servidor do modelo e escuta o EventBus para ver as duas camadas em ação.
Do zero: O loop do nível 2, com uma função por camada: callModel tenta o servidor de
novo com backoff, runTool transforma cada falha de ferramenta em texto que o modelo consegue ler, e
todo o resto fica livre para lançar. Por cima disso, um pequeno contador desiste quando o modelo continua batendo
na mesma parede.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const room = { lockOiled: false, doorOpen: false }
const use = tool({
name: 'use',
description: 'Use an inventory item with something in the room.',
schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
execute: async ({ item, target }) => {
if (item === 'oil_can' && target === 'lock') {
room.lockOiled = true
return 'You oil the lock. It looks like it might turn now.'
}
if (item === 'key' && target === 'door') {
// Just throw. astorlm catches it and sends the message back to the model, marked is_error.
if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
room.doorOpen = true
return 'Click! The key turns and the door swings open.'
}
throw new Error(`Nothing happens when you use ${item} with ${target}.`)
},
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [lookAt, use],
// Off by default: without it, a single 429 ends the run.
retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})
// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
if (event.type === 'provider_retry') {
console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
}
})
try {
const last = await agent.run('Open the temple door.')
console.log(last.content)
} catch (err) {
// Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
console.error('The run failed:', err)
}
// Errors in an agent loop, from scratch. Plain fetch, no SDK.
// The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
const room = { lockOiled: false, doorOpen: false }
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { look_at: lookAt, use }
const toolSchemas = [/* one JSON Schema per tool */]
async function use({ item, target }: Record<string, string>): Promise<string> {
if (item === 'oil_can' && target === 'lock') {
room.lockOiled = true
return 'You oil the lock. It looks like it might turn now.'
}
if (item === 'key' && target === 'door') {
// Say what went wrong and what to try instead: this text is all the model will get.
if (!room.lockOiled) throw new Error('the lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
room.doorOpen = true
return 'Click! The key turns and the door swings open.'
}
throw new Error(`nothing happens when you use ${item} with ${target}.`)
}
// Layer 1, the model's server. Retry only what is likely to pass by itself.
async function callModel(messages: Message[], maxAttempts = 3) {
for (let attempt = 1; ; attempt++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
}).catch(() => null) // null: no response at all, the network failed
if (res?.ok) return (await res.json()).choices[0]
// 429 (too many requests), 5xx (server trouble) and network errors tend to pass.
// A 400 or a 401 won't: retrying just sends the same broken request again.
const transient = res === null || res.status === 429 || res.status >= 500
if (!transient || attempt === maxAttempts) throw new Error(`Model call failed: ${res?.status ?? 'network error'}`)
// Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
await new Promise((resolve) => setTimeout(resolve, Math.random() * 1000 * 2 ** (attempt - 1)))
}
}
// Layer 2, the tools. Whatever goes wrong becomes a result the model can read.
async function runTool(call: ToolCall): Promise<{ output: string; isError: boolean }> {
const run = tools[call.function.name]
if (!run) return { output: `Unknown tool: ${call.function.name}. Tools: ${Object.keys(tools).join(', ')}`, isError: true }
try {
const args = JSON.parse(call.function.arguments) // models do send broken JSON now and then
return { output: await run(args), isError: false }
} catch (err) {
return { output: `Error: ${err instanceof Error ? err.message : err}`, isError: true }
}
}
export async function runAgent(prompt: string, maxTurns = 10, maxErrorsInARow = 3): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
let errorsInARow = 0
for (let turn = 1; turn <= maxTurns; turn++) {
const choice = await callModel(messages)
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const { output, isError } = await runTool(call)
// Chat Completions has no is_error field: the text itself has to say it failed.
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
errorsInARow = isError ? errorsInARow + 1 : 0
}
// A model stuck on the same wrong move rarely gets out by itself.
if (errorsInARow >= maxErrorsInARow) throw new Error(`Gave up after ${errorsInARow} tool errors in a row`)
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// Layer 3 is everything else: a bug in this file, an abort. Nothing catches it here, on purpose.
await runAgent('Open the temple door.')
# Errors in an agent loop, from scratch. Standard library only, no SDK.
# The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
import json
import random
import time
import urllib.error
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
room = {"lock_oiled": False, "door_open": False}
def use(item, target):
if item == "oil_can" and target == "lock":
room["lock_oiled"] = True
return "You oil the lock. It looks like it might turn now."
if item == "key" and target == "door":
# Say what went wrong and what to try instead: this text is all the model will get.
if not room["lock_oiled"]:
raise ValueError("the lock is rusted shut and the key won't turn. Oil it first: use oil_can with lock.")
room["door_open"] = True
return "Click! The key turns and the door swings open."
raise ValueError(f"nothing happens when you use {item} with {target}.")
TOOLS = {"look_at": look_at, "use": use}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
def call_model(messages, max_attempts=3):
"""Layer 1, the model's server. Retry only what is likely to pass by itself."""
for attempt in range(1, max_attempts + 1):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
try:
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
except urllib.error.HTTPError as err:
# 429 (too many requests) and 5xx (server trouble) tend to pass.
# A 400 or a 401 won't: retrying just sends the same broken request again.
if (err.code != 429 and err.code < 500) or attempt == max_attempts:
raise
except urllib.error.URLError:
# No response at all: the network failed.
if attempt == max_attempts:
raise
# Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
time.sleep(random.random() * 2 ** (attempt - 1))
def run_tool(call):
"""Layer 2, the tools. Whatever goes wrong becomes a result the model can read."""
name = call["function"]["name"]
run = TOOLS.get(name)
if run is None:
return f"Unknown tool: {name}. Tools: {', '.join(TOOLS)}", True
try:
args = json.loads(call["function"]["arguments"]) # models do send broken JSON now and then
return run(**args), False
except Exception as err:
return f"Error: {err}", True
def run_agent(prompt, max_turns=10, max_errors_in_a_row=3):
messages = [{"role": "user", "content": prompt}]
errors_in_a_row = 0
for _ in range(max_turns):
choice = call_model(messages)
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
output, is_error = run_tool(call)
# Chat Completions has no is_error field: the text itself has to say it failed.
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
errors_in_a_row = errors_in_a_row + 1 if is_error else 0
# A model stuck on the same wrong move rarely gets out by itself.
if errors_in_a_row >= max_errors_in_a_row:
raise RuntimeError(f"Gave up after {errors_in_a_row} tool errors in a row")
raise RuntimeError(f"No answer after {max_turns} turns")
# Layer 3 is everything else: a bug in this file, a Ctrl+C. Nothing catches it here, on purpose.
run_agent("Open the temple door.")
O que observar
- Escreva os erros para o modelo. Diga o que estava errado e o que fazer no lugar: "a fechadura está emperrada de ferrugem. Lubrifique primeiro: use oil_can com lock". Com "Error 400", você recebe a mesma chamada de novo.
- Não vaze o que o modelo não deveria ver. Stack traces, caminhos de arquivo e strings de conexão vão para os seus logs. O modelo recebe uma frase clara.
-
Erros também podem entrar em loop. Um modelo pode tentar o mesmo movimento errado até o
maxTurnsacabar. Pare depois de alguns erros seguidos, ou capture isso com um hook (próximo nível). - Tente de novo com limite e espera aleatória. Poucas tentativas, um teto que dobra a cada vez e um pouco de aleatoriedade (jitter) para que cem clientes não voltem todos no mesmo segundo.
- Cuidado ao repetir ferramentas que mudam coisas. Se uma chamada que movimenta dinheiro ou reserva um quarto deu timeout, ela pode ter passado mesmo assim. Confira antes de fazer duas vezes.