> Nivel 5 de Agent Harness Patterns, un recorrido de patrones sobre cómo funcionan los agentes de IA. Versión web: https://harnesspatterns.dev/es/patterns/errors-in-the-loop · Todos los patrones (en inglés): https://harnesspatterns.dev/llms.txt

# Errores en el bucle

Las herramientas fallan. Los modelos mandan datos incorrectos. Los servidores se caen por un segundo. Un bucle de agente tiene que decidir, para cada falla, quién se encarga: el modelo, tu código o nadie.

## El problema

Un agente juega una aventura gráfica. Sus herramientas son los verbos de la parte de abajo de la pantalla, `look_at` y `use`, y la tarea es "abre la puerta del templo". El modelo hace lo obvio: usar la llave con la puerta. Pero la cerradura está oxidada, y la herramienta lanza una excepción.

Las viejas aventuras gráficas se dividían en dos escuelas ante momentos así. En algunas, un movimiento equivocado te mataba: game over, de vuelta a tu última partida guardada. En otras no podías morir; el juego solo te decía por qué no funcionaba, y probabas otra cosa. Un bucle de agente también tiene que elegir escuela.

Si el bucle no atrapa el error, gana Krash: la excepción sube a través del bucle y `run()` la lanza. Se pierde la ejecución, y el error decía exactamente qué hacer: "aceítala primero". El único lector que podía usar ese mensaje nunca lo recibió. Atrápalo, devuélvelo como resultado de herramienta, y el modelo lo lee y lo vuelve a intentar. Eso cuesta un turno. No atraparlo cuesta toda la ejecución.

## Tres tipos de falla

- El modelo pidió mal
   Entrada inválida, JSON roto, una herramienta que no existe, una llave que no gira.
   **El modelo.** El error vuelve como resultado de herramienta marcado como error. El modelo lo lee y prueba otra cosa en el siguiente turno.
- Falló el servidor del modelo
   429 (demasiados pedidos), 5xx (problemas del servidor), una conexión que se corta.
   **Tu código, reintentando.** Espera un poco y vuelve a llamar, un poco más cada vez. El modelo nunca se entera: no hubo ningún turno que leer.
- Nadie puede arreglarlo
   Un 401 (API key incorrecta), un bug en tu código, el usuario canceló.
   **Nadie.** Deja que lance la excepción. Reintentar no va a servir, y ocultárselo al modelo solo lo hace adivinar.

El truco es no mezclarlas. Reintentar un 400 vuelve a mandar el mismo pedido roto. Mostrarle un 503 al modelo gasta un turno en algo que no puede arreglar. Y atrapar un bug de tu propio código solo lo esconde.

> Hoy en astorlm
>
> Los errores de herramientas se manejan por ti: el ToolRegistry atrapa una herramienta desconocida, una entrada que no pasa el esquema de Zod y cualquier cosa que lance `execute`, y la devuelve con `is_error: true`. Reintentar el modelo viene **apagado por defecto**: sin la opción `retry`, un solo 503 termina la ejecución con `session_end: error`. Y un error de esquema le llega al modelo como el ZodError crudo, un volcado JSON de cada problema. Funciona, pero una frase corta escrita por ti se leería mejor.

## El código

**Con astorlm:** Las herramientas simplemente lanzan excepciones, y astorlm se encarga de atraparlas. Tú activas `retry` para el servidor del modelo y escuchas el EventBus para ver las dos capas en acción.

**Desde cero:** El bucle del nivel 2, con una función por capa: `callModel` reintenta contra el servidor con backoff, `runTool` convierte cada falla de herramienta en texto que el modelo puede leer, y todo lo demás se deja lanzar. Encima de eso, un pequeño contador se rinde cuando el modelo sigue chocando contra la misma pared.

**Con astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const room = { lockOiled: false, doorOpen: false }

const use = tool({
  name: 'use',
  description: 'Use an inventory item with something in the room.',
  schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
  execute: async ({ item, target }) => {
    if (item === 'oil_can' && target === 'lock') {
      room.lockOiled = true
      return 'You oil the lock. It looks like it might turn now.'
    }
    if (item === 'key' && target === 'door') {
      // Just throw. astorlm catches it and sends the message back to the model, marked is_error.
      if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
      room.doorOpen = true
      return 'Click! The key turns and the door swings open.'
    }
    throw new Error(`Nothing happens when you use ${item} with ${target}.`)
  },
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [lookAt, use],
  // Off by default: without it, a single 429 ends the run.
  retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
  maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})

// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
  if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
  if (event.type === 'provider_retry') {
    console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
  }
})

try {
  const last = await agent.run('Open the temple door.')
  console.log(last.content)
} catch (err) {
  // Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
  console.error('The run failed:', err)
}
```

**TypeScript**

```ts
// Errors in an agent loop, from scratch. Plain fetch, no SDK.
// The agent plays an adventure game: its tools are the verbs LOOK AT and USE.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

const room = { lockOiled: false, doorOpen: false }

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { look_at: lookAt, use }
const toolSchemas = [/* one JSON Schema per tool */]

async function use({ item, target }: Record<string, string>): Promise<string> {
  if (item === 'oil_can' && target === 'lock') {
    room.lockOiled = true
    return 'You oil the lock. It looks like it might turn now.'
  }
  if (item === 'key' && target === 'door') {
    // Say what went wrong and what to try instead: this text is all the model will get.
    if (!room.lockOiled) throw new Error('the lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
    room.doorOpen = true
    return 'Click! The key turns and the door swings open.'
  }
  throw new Error(`nothing happens when you use ${item} with ${target}.`)
}

// Layer 1, the model's server. Retry only what is likely to pass by itself.
async function callModel(messages: Message[], maxAttempts = 3) {
  for (let attempt = 1; ; attempt++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    }).catch(() => null) // null: no response at all, the network failed
    if (res?.ok) return (await res.json()).choices[0]

    // 429 (too many requests), 5xx (server trouble) and network errors tend to pass.
    // A 400 or a 401 won't: retrying just sends the same broken request again.
    const transient = res === null || res.status === 429 || res.status >= 500
    if (!transient || attempt === maxAttempts) throw new Error(`Model call failed: ${res?.status ?? 'network error'}`)
    // Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
    await new Promise((resolve) => setTimeout(resolve, Math.random() * 1000 * 2 ** (attempt - 1)))
  }
}

// Layer 2, the tools. Whatever goes wrong becomes a result the model can read.
async function runTool(call: ToolCall): Promise<{ output: string; isError: boolean }> {
  const run = tools[call.function.name]
  if (!run) return { output: `Unknown tool: ${call.function.name}. Tools: ${Object.keys(tools).join(', ')}`, isError: true }
  try {
    const args = JSON.parse(call.function.arguments) // models do send broken JSON now and then
    return { output: await run(args), isError: false }
  } catch (err) {
    return { output: `Error: ${err instanceof Error ? err.message : err}`, isError: true }
  }
}

export async function runAgent(prompt: string, maxTurns = 10, maxErrorsInARow = 3): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]
  let errorsInARow = 0

  for (let turn = 1; turn <= maxTurns; turn++) {
    const choice = await callModel(messages)
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const { output, isError } = await runTool(call)
      // Chat Completions has no is_error field: the text itself has to say it failed.
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
      errorsInARow = isError ? errorsInARow + 1 : 0
    }
    // A model stuck on the same wrong move rarely gets out by itself.
    if (errorsInARow >= maxErrorsInARow) throw new Error(`Gave up after ${errorsInARow} tool errors in a row`)
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
// Layer 3 is everything else: a bug in this file, an abort. Nothing catches it here, on purpose.

await runAgent('Open the temple door.')
```

**Python**

```python
# Errors in an agent loop, from scratch. Standard library only, no SDK.
# The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
import json
import random
import time
import urllib.error
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

room = {"lock_oiled": False, "door_open": False}

def use(item, target):
    if item == "oil_can" and target == "lock":
        room["lock_oiled"] = True
        return "You oil the lock. It looks like it might turn now."
    if item == "key" and target == "door":
        # Say what went wrong and what to try instead: this text is all the model will get.
        if not room["lock_oiled"]:
            raise ValueError("the lock is rusted shut and the key won't turn. Oil it first: use oil_can with lock.")
        room["door_open"] = True
        return "Click! The key turns and the door swings open."
    raise ValueError(f"nothing happens when you use {item} with {target}.")

TOOLS = {"look_at": look_at, "use": use}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

def call_model(messages, max_attempts=3):
    """Layer 1, the model's server. Retry only what is likely to pass by itself."""
    for attempt in range(1, max_attempts + 1):
        request = urllib.request.Request(
            f"{LLM['base_url']}/chat/completions",
            data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
            headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
        )
        try:
            with urllib.request.urlopen(request) as response:
                return json.load(response)["choices"][0]
        except urllib.error.HTTPError as err:
            # 429 (too many requests) and 5xx (server trouble) tend to pass.
            # A 400 or a 401 won't: retrying just sends the same broken request again.
            if (err.code != 429 and err.code < 500) or attempt == max_attempts:
                raise
        except urllib.error.URLError:
            # No response at all: the network failed.
            if attempt == max_attempts:
                raise
        # Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
        time.sleep(random.random() * 2 ** (attempt - 1))

def run_tool(call):
    """Layer 2, the tools. Whatever goes wrong becomes a result the model can read."""
    name = call["function"]["name"]
    run = TOOLS.get(name)
    if run is None:
        return f"Unknown tool: {name}. Tools: {', '.join(TOOLS)}", True
    try:
        args = json.loads(call["function"]["arguments"])  # models do send broken JSON now and then
        return run(**args), False
    except Exception as err:
        return f"Error: {err}", True

def run_agent(prompt, max_turns=10, max_errors_in_a_row=3):
    messages = [{"role": "user", "content": prompt}]
    errors_in_a_row = 0

    for _ in range(max_turns):
        choice = call_model(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output, is_error = run_tool(call)
            # Chat Completions has no is_error field: the text itself has to say it failed.
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
            errors_in_a_row = errors_in_a_row + 1 if is_error else 0

        # A model stuck on the same wrong move rarely gets out by itself.
        if errors_in_a_row >= max_errors_in_a_row:
            raise RuntimeError(f"Gave up after {errors_in_a_row} tool errors in a row")

    raise RuntimeError(f"No answer after {max_turns} turns")
# Layer 3 is everything else: a bug in this file, a Ctrl+C. Nothing catches it here, on purpose.

run_agent("Open the temple door.")
```

## Qué vigilar

- **Escribe los errores para el modelo.** Di qué estuvo mal y qué hacer en cambio: "la cerradura está trabada por el óxido. Aceítala primero: usa oil_can con lock". Con "Error 400" vas a recibir la misma llamada otra vez.
- **No filtres lo que el modelo no debería ver.** Los stack traces, las rutas de archivos y las cadenas de conexión van a tus logs. El modelo recibe una frase clara.
- **Los errores también pueden entrar en bucle.** Un modelo puede intentar el mismo movimiento equivocado hasta que se agote `maxTurns`. Detente después de unos cuantos errores seguidos, o atrápalo con un hook (próximo nivel).
- **Reintenta con un límite y una espera aleatoria.** Unos pocos intentos, un techo que se duplica cada vez y algo de aleatoriedad (jitter) para que cien clientes no vuelvan todos en el mismo segundo.
- **Cuidado al reintentar herramientas que cambian cosas.** Si una llamada que mueve dinero o reserva una habitación dio timeout, puede que se haya hecho igual. Verifica antes de hacerla dos veces.

## Patrones relacionados

- [3 · Diseñar una herramienta](https://harnesspatterns.dev/es/patterns/designing-a-tool.md)
- [4 · Cuándo parar](https://harnesspatterns.dev/es/patterns/when-to-stop.md)
- [6 · Hooks](https://harnesspatterns.dev/es/patterns/hooks.md)
- [11 · Observabilidad y evaluaciones](https://harnesspatterns.dev/es/patterns/observability.md)
