> Level 5 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/errors-in-the-loop · All patterns: https://harnesspatterns.dev/llms.txt

# Errors in the loop

Tools fail. Models send bad input. Servers go down for a second. An agent loop has to decide, for each failure, who deals with it: the model, your code, or nobody.

## The problem

An agent plays an adventure game. Its tools are the verbs at the bottom of the screen, `look_at` and `use`, and the task is "open the temple door". The model does the obvious thing: use the key with the door. But the lock is rusted, and the tool throws.

Old adventure games split into two schools over moments like this. In some, a wrong move killed you: game over, back to your last save. In others you couldn't die; the game just told you why it didn't work, and you tried something else. An agent loop has to pick a school too.

If the loop doesn't catch the error, Krash wins: the exception goes up through the loop and `run()` throws. The run is lost, and the error said exactly what to do: "oil it first". The one reader who could use that message never got it. Catch it, send it back as a tool result, and the model reads it and tries again. That costs a turn. Not catching it costs the whole run.

## Three kinds of failure

- The model asked wrong
   Invalid input, broken JSON, a tool that doesn’t exist, a key that won’t turn.
   **The model.** The error goes back as a tool result marked as an error. The model reads it and tries something else on the next turn.
- The model’s server failed
   429 (too many requests), 5xx (server trouble), a dropped connection.
   **Your code, retrying.** Wait a little and call again, a bit longer each time. The model never finds out: there was no turn to read.
- Nobody can fix it
   A 401 (bad API key), a bug in your code, the user cancelled.
   **Nobody.** Let it throw. Retrying won’t help, and hiding it from the model only makes the model guess.

The trick is not to mix them up. Retrying a 400 sends the same broken request again. Showing a 503 to the model spends a turn on something it can't fix. And catching a bug in your own code just hides it.

> In astorlm today
>
> Tool errors are handled for you: the ToolRegistry catches an unknown tool, input that fails the Zod schema, and anything `execute` throws, and sends it back with `is_error: true`. Retrying the model is **off by default**: without the `retry` option, a single 503 ends the run with `session_end: error`. And a schema error reaches the model as the raw ZodError, a JSON dump of every issue. It works, but a short sentence of your own would read better.

## The code

**With astorlm:** Tools just throw, and astorlm does the catching. You turn on `retry` for the model's server, and listen to the EventBus to see both layers at work.

**From scratch:** The loop from level 2, with one function per layer: `callModel` retries the server with backoff, `runTool` turns every tool failure into text the model can read, and anything else is left to throw. On top of that, a small counter gives up when the model keeps hitting the same wall.

**With astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const room = { lockOiled: false, doorOpen: false }

const use = tool({
  name: 'use',
  description: 'Use an inventory item with something in the room.',
  schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
  execute: async ({ item, target }) => {
    if (item === 'oil_can' && target === 'lock') {
      room.lockOiled = true
      return 'You oil the lock. It looks like it might turn now.'
    }
    if (item === 'key' && target === 'door') {
      // Just throw. astorlm catches it and sends the message back to the model, marked is_error.
      if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
      room.doorOpen = true
      return 'Click! The key turns and the door swings open.'
    }
    throw new Error(`Nothing happens when you use ${item} with ${target}.`)
  },
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [lookAt, use],
  // Off by default: without it, a single 429 ends the run.
  retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
  maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})

// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
  if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
  if (event.type === 'provider_retry') {
    console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
  }
})

try {
  const last = await agent.run('Open the temple door.')
  console.log(last.content)
} catch (err) {
  // Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
  console.error('The run failed:', err)
}
```

**TypeScript**

```ts
// Errors in an agent loop, from scratch. Plain fetch, no SDK.
// The agent plays an adventure game: its tools are the verbs LOOK AT and USE.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

const room = { lockOiled: false, doorOpen: false }

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { look_at: lookAt, use }
const toolSchemas = [/* one JSON Schema per tool */]

async function use({ item, target }: Record<string, string>): Promise<string> {
  if (item === 'oil_can' && target === 'lock') {
    room.lockOiled = true
    return 'You oil the lock. It looks like it might turn now.'
  }
  if (item === 'key' && target === 'door') {
    // Say what went wrong and what to try instead: this text is all the model will get.
    if (!room.lockOiled) throw new Error('the lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
    room.doorOpen = true
    return 'Click! The key turns and the door swings open.'
  }
  throw new Error(`nothing happens when you use ${item} with ${target}.`)
}

// Layer 1, the model's server. Retry only what is likely to pass by itself.
async function callModel(messages: Message[], maxAttempts = 3) {
  for (let attempt = 1; ; attempt++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    }).catch(() => null) // null: no response at all, the network failed
    if (res?.ok) return (await res.json()).choices[0]

    // 429 (too many requests), 5xx (server trouble) and network errors tend to pass.
    // A 400 or a 401 won't: retrying just sends the same broken request again.
    const transient = res === null || res.status === 429 || res.status >= 500
    if (!transient || attempt === maxAttempts) throw new Error(`Model call failed: ${res?.status ?? 'network error'}`)
    // Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
    await new Promise((resolve) => setTimeout(resolve, Math.random() * 1000 * 2 ** (attempt - 1)))
  }
}

// Layer 2, the tools. Whatever goes wrong becomes a result the model can read.
async function runTool(call: ToolCall): Promise<{ output: string; isError: boolean }> {
  const run = tools[call.function.name]
  if (!run) return { output: `Unknown tool: ${call.function.name}. Tools: ${Object.keys(tools).join(', ')}`, isError: true }
  try {
    const args = JSON.parse(call.function.arguments) // models do send broken JSON now and then
    return { output: await run(args), isError: false }
  } catch (err) {
    return { output: `Error: ${err instanceof Error ? err.message : err}`, isError: true }
  }
}

export async function runAgent(prompt: string, maxTurns = 10, maxErrorsInARow = 3): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]
  let errorsInARow = 0

  for (let turn = 1; turn <= maxTurns; turn++) {
    const choice = await callModel(messages)
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const { output, isError } = await runTool(call)
      // Chat Completions has no is_error field: the text itself has to say it failed.
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
      errorsInARow = isError ? errorsInARow + 1 : 0
    }
    // A model stuck on the same wrong move rarely gets out by itself.
    if (errorsInARow >= maxErrorsInARow) throw new Error(`Gave up after ${errorsInARow} tool errors in a row`)
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
// Layer 3 is everything else: a bug in this file, an abort. Nothing catches it here, on purpose.

await runAgent('Open the temple door.')
```

**Python**

```python
# Errors in an agent loop, from scratch. Standard library only, no SDK.
# The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
import json
import random
import time
import urllib.error
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

room = {"lock_oiled": False, "door_open": False}

def use(item, target):
    if item == "oil_can" and target == "lock":
        room["lock_oiled"] = True
        return "You oil the lock. It looks like it might turn now."
    if item == "key" and target == "door":
        # Say what went wrong and what to try instead: this text is all the model will get.
        if not room["lock_oiled"]:
            raise ValueError("the lock is rusted shut and the key won't turn. Oil it first: use oil_can with lock.")
        room["door_open"] = True
        return "Click! The key turns and the door swings open."
    raise ValueError(f"nothing happens when you use {item} with {target}.")

TOOLS = {"look_at": look_at, "use": use}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

def call_model(messages, max_attempts=3):
    """Layer 1, the model's server. Retry only what is likely to pass by itself."""
    for attempt in range(1, max_attempts + 1):
        request = urllib.request.Request(
            f"{LLM['base_url']}/chat/completions",
            data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
            headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
        )
        try:
            with urllib.request.urlopen(request) as response:
                return json.load(response)["choices"][0]
        except urllib.error.HTTPError as err:
            # 429 (too many requests) and 5xx (server trouble) tend to pass.
            # A 400 or a 401 won't: retrying just sends the same broken request again.
            if (err.code != 429 and err.code < 500) or attempt == max_attempts:
                raise
        except urllib.error.URLError:
            # No response at all: the network failed.
            if attempt == max_attempts:
                raise
        # Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
        time.sleep(random.random() * 2 ** (attempt - 1))

def run_tool(call):
    """Layer 2, the tools. Whatever goes wrong becomes a result the model can read."""
    name = call["function"]["name"]
    run = TOOLS.get(name)
    if run is None:
        return f"Unknown tool: {name}. Tools: {', '.join(TOOLS)}", True
    try:
        args = json.loads(call["function"]["arguments"])  # models do send broken JSON now and then
        return run(**args), False
    except Exception as err:
        return f"Error: {err}", True

def run_agent(prompt, max_turns=10, max_errors_in_a_row=3):
    messages = [{"role": "user", "content": prompt}]
    errors_in_a_row = 0

    for _ in range(max_turns):
        choice = call_model(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output, is_error = run_tool(call)
            # Chat Completions has no is_error field: the text itself has to say it failed.
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
            errors_in_a_row = errors_in_a_row + 1 if is_error else 0

        # A model stuck on the same wrong move rarely gets out by itself.
        if errors_in_a_row >= max_errors_in_a_row:
            raise RuntimeError(f"Gave up after {errors_in_a_row} tool errors in a row")

    raise RuntimeError(f"No answer after {max_turns} turns")
# Layer 3 is everything else: a bug in this file, a Ctrl+C. Nothing catches it here, on purpose.

run_agent("Open the temple door.")
```

## What to watch

- **Write errors for the model.** Say what was wrong and what to do instead: "the lock is rusted shut. Oil it first: use oil_can with lock". "Error 400" gets you the same call again.
- **Don't leak what the model shouldn't see.** Stack traces, file paths and connection strings go in your logs. The model gets one clear sentence.
- **Errors can loop too.** A model can try the same wrong move until `maxTurns` runs out. Stop after a few errors in a row, or catch it with a hook (next level).
- **Retry with a limit and a random wait.** A few attempts, a ceiling that doubles each time, and some randomness (jitter) so a hundred clients don't all come back at the same second.
- **Be careful retrying tools that change things.** If a call that moves money or books a room timed out, it may have gone through anyway. Check before doing it twice.

## Related patterns

- [3 · Designing a tool](https://harnesspatterns.dev/patterns/designing-a-tool.md)
- [4 · When to stop](https://harnesspatterns.dev/patterns/when-to-stop.md)
- [6 · Hooks](https://harnesspatterns.dev/patterns/hooks.md)
- [11 · Observability and evals](https://harnesspatterns.dev/patterns/observability.md)
