> Niveau 5 de Agent Harness Patterns, un parcours de patterns sur le fonctionnement des agents d'IA. Version web : https://harnesspatterns.dev/fr/patterns/errors-in-the-loop · Tous les patterns (en anglais) : https://harnesspatterns.dev/llms.txt

# Les erreurs dans la boucle

Les outils échouent. Les modèles envoient de mauvaises entrées. Les serveurs tombent une seconde. Une boucle d'agent doit décider, pour chaque échec, qui s'en occupe : le modèle, votre code, ou personne.

## Le problème

Un agent joue à un jeu d'aventure. Ses outils sont les verbes en bas de l'écran, `look_at` et `use`, et la tâche est « ouvre la porte du temple ». Le modèle fait l'évidence : utiliser la clé sur la porte. Mais la serrure est rouillée, et l'outil lève une exception.

Face à ce genre de moment, les vieux jeux d'aventure se divisaient en deux écoles. Dans certains, un mauvais coup vous tuait : game over, retour à votre dernière sauvegarde. Dans d'autres, on ne pouvait pas mourir ; le jeu vous disait simplement pourquoi ça ne marchait pas, et vous essayiez autre chose. Une boucle d'agent doit, elle aussi, choisir son école.

Si la boucle n'attrape pas l'erreur, Krash gagne : l'exception remonte à travers la boucle et `run()` la lève. L'exécution est perdue, alors que l'erreur disait exactement quoi faire : « huile-la d'abord ». Le seul lecteur qui pouvait se servir de ce message ne l'a jamais reçu. Attrapez-la, renvoyez-la comme résultat d'outil, et le modèle la lit et réessaie. Cela coûte un tour. Ne pas l'attraper coûte toute l'exécution.

## Trois sortes d'échecs

- Le modèle a mal demandé
   Entrée invalide, JSON cassé, un outil qui n'existe pas, une clé qui ne tourne pas.
   **Le modèle.** L'erreur revient comme un résultat d'outil marqué en erreur. Le modèle la lit et essaie autre chose au tour suivant.
- Le serveur du modèle a échoué
   429 (trop de requêtes), 5xx (problème côté serveur), une connexion coupée.
   **Votre code, en réessayant.** Attendez un peu et rappelez, un peu plus longtemps à chaque fois. Le modèle n'en sait jamais rien : il n'y a eu aucun tour à lire.
- Personne ne peut corriger
   Un 401 (mauvaise clé d'API), un bug dans votre code, l'utilisateur a annulé.
   **Personne.** Laissez l'exception remonter. Réessayer n'y changera rien, et le cacher au modèle ne fait que le pousser à deviner.

Toute l'astuce est de ne pas les mélanger. Réessayer un 400 renvoie la même requête cassée. Montrer un 503 au modèle gaspille un tour sur quelque chose qu'il ne peut pas réparer. Et attraper un bug de votre propre code ne fait que le cacher.

> Aujourd'hui dans astorlm
>
> Les erreurs d'outils sont gérées pour vous : le ToolRegistry attrape un outil inconnu, une entrée qui ne passe pas le schéma Zod et tout ce que lève `execute`, et le renvoie avec `is_error: true`. Réessayer le modèle est **désactivé par défaut** : sans l'option `retry`, un seul 503 met fin à l'exécution avec `session_end: error`. Et une erreur de schéma arrive au modèle sous forme de ZodError brut, un dump JSON de chaque problème. Ça marche, mais une courte phrase écrite par vous serait plus lisible.

## Le code

**Avec astorlm :** Les outils se contentent de lever des exceptions, et astorlm se charge de les attraper. Vous activez `retry` pour le serveur du modèle, et vous écoutez l'EventBus pour voir les deux couches à l'œuvre.

**À partir de zéro :** La boucle du niveau 2, avec une fonction par couche : `callModel` réessaie le serveur avec backoff, `runTool` transforme chaque échec d'outil en texte lisible par le modèle, et tout le reste est laissé libre de lever. Par-dessus, un petit compteur abandonne quand le modèle se heurte sans cesse au même mur.

**Avec astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const room = { lockOiled: false, doorOpen: false }

const use = tool({
  name: 'use',
  description: 'Use an inventory item with something in the room.',
  schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
  execute: async ({ item, target }) => {
    if (item === 'oil_can' && target === 'lock') {
      room.lockOiled = true
      return 'You oil the lock. It looks like it might turn now.'
    }
    if (item === 'key' && target === 'door') {
      // Just throw. astorlm catches it and sends the message back to the model, marked is_error.
      if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
      room.doorOpen = true
      return 'Click! The key turns and the door swings open.'
    }
    throw new Error(`Nothing happens when you use ${item} with ${target}.`)
  },
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [lookAt, use],
  // Off by default: without it, a single 429 ends the run.
  retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
  maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})

// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
  if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
  if (event.type === 'provider_retry') {
    console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
  }
})

try {
  const last = await agent.run('Open the temple door.')
  console.log(last.content)
} catch (err) {
  // Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
  console.error('The run failed:', err)
}
```

**TypeScript**

```ts
// Errors in an agent loop, from scratch. Plain fetch, no SDK.
// The agent plays an adventure game: its tools are the verbs LOOK AT and USE.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

const room = { lockOiled: false, doorOpen: false }

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { look_at: lookAt, use }
const toolSchemas = [/* one JSON Schema per tool */]

async function use({ item, target }: Record<string, string>): Promise<string> {
  if (item === 'oil_can' && target === 'lock') {
    room.lockOiled = true
    return 'You oil the lock. It looks like it might turn now.'
  }
  if (item === 'key' && target === 'door') {
    // Say what went wrong and what to try instead: this text is all the model will get.
    if (!room.lockOiled) throw new Error('the lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
    room.doorOpen = true
    return 'Click! The key turns and the door swings open.'
  }
  throw new Error(`nothing happens when you use ${item} with ${target}.`)
}

// Layer 1, the model's server. Retry only what is likely to pass by itself.
async function callModel(messages: Message[], maxAttempts = 3) {
  for (let attempt = 1; ; attempt++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    }).catch(() => null) // null: no response at all, the network failed
    if (res?.ok) return (await res.json()).choices[0]

    // 429 (too many requests), 5xx (server trouble) and network errors tend to pass.
    // A 400 or a 401 won't: retrying just sends the same broken request again.
    const transient = res === null || res.status === 429 || res.status >= 500
    if (!transient || attempt === maxAttempts) throw new Error(`Model call failed: ${res?.status ?? 'network error'}`)
    // Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
    await new Promise((resolve) => setTimeout(resolve, Math.random() * 1000 * 2 ** (attempt - 1)))
  }
}

// Layer 2, the tools. Whatever goes wrong becomes a result the model can read.
async function runTool(call: ToolCall): Promise<{ output: string; isError: boolean }> {
  const run = tools[call.function.name]
  if (!run) return { output: `Unknown tool: ${call.function.name}. Tools: ${Object.keys(tools).join(', ')}`, isError: true }
  try {
    const args = JSON.parse(call.function.arguments) // models do send broken JSON now and then
    return { output: await run(args), isError: false }
  } catch (err) {
    return { output: `Error: ${err instanceof Error ? err.message : err}`, isError: true }
  }
}

export async function runAgent(prompt: string, maxTurns = 10, maxErrorsInARow = 3): Promise<string> {
  const messages: Message[] = [{ role: 'user', content: prompt }]
  let errorsInARow = 0

  for (let turn = 1; turn <= maxTurns; turn++) {
    const choice = await callModel(messages)
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const { output, isError } = await runTool(call)
      // Chat Completions has no is_error field: the text itself has to say it failed.
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
      errorsInARow = isError ? errorsInARow + 1 : 0
    }
    // A model stuck on the same wrong move rarely gets out by itself.
    if (errorsInARow >= maxErrorsInARow) throw new Error(`Gave up after ${errorsInARow} tool errors in a row`)
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
// Layer 3 is everything else: a bug in this file, an abort. Nothing catches it here, on purpose.

await runAgent('Open the temple door.')
```

**Python**

```python
# Errors in an agent loop, from scratch. Standard library only, no SDK.
# The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
import json
import random
import time
import urllib.error
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

room = {"lock_oiled": False, "door_open": False}

def use(item, target):
    if item == "oil_can" and target == "lock":
        room["lock_oiled"] = True
        return "You oil the lock. It looks like it might turn now."
    if item == "key" and target == "door":
        # Say what went wrong and what to try instead: this text is all the model will get.
        if not room["lock_oiled"]:
            raise ValueError("the lock is rusted shut and the key won't turn. Oil it first: use oil_can with lock.")
        room["door_open"] = True
        return "Click! The key turns and the door swings open."
    raise ValueError(f"nothing happens when you use {item} with {target}.")

TOOLS = {"look_at": look_at, "use": use}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

def call_model(messages, max_attempts=3):
    """Layer 1, the model's server. Retry only what is likely to pass by itself."""
    for attempt in range(1, max_attempts + 1):
        request = urllib.request.Request(
            f"{LLM['base_url']}/chat/completions",
            data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
            headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
        )
        try:
            with urllib.request.urlopen(request) as response:
                return json.load(response)["choices"][0]
        except urllib.error.HTTPError as err:
            # 429 (too many requests) and 5xx (server trouble) tend to pass.
            # A 400 or a 401 won't: retrying just sends the same broken request again.
            if (err.code != 429 and err.code < 500) or attempt == max_attempts:
                raise
        except urllib.error.URLError:
            # No response at all: the network failed.
            if attempt == max_attempts:
                raise
        # Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
        time.sleep(random.random() * 2 ** (attempt - 1))

def run_tool(call):
    """Layer 2, the tools. Whatever goes wrong becomes a result the model can read."""
    name = call["function"]["name"]
    run = TOOLS.get(name)
    if run is None:
        return f"Unknown tool: {name}. Tools: {', '.join(TOOLS)}", True
    try:
        args = json.loads(call["function"]["arguments"])  # models do send broken JSON now and then
        return run(**args), False
    except Exception as err:
        return f"Error: {err}", True

def run_agent(prompt, max_turns=10, max_errors_in_a_row=3):
    messages = [{"role": "user", "content": prompt}]
    errors_in_a_row = 0

    for _ in range(max_turns):
        choice = call_model(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output, is_error = run_tool(call)
            # Chat Completions has no is_error field: the text itself has to say it failed.
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
            errors_in_a_row = errors_in_a_row + 1 if is_error else 0

        # A model stuck on the same wrong move rarely gets out by itself.
        if errors_in_a_row >= max_errors_in_a_row:
            raise RuntimeError(f"Gave up after {errors_in_a_row} tool errors in a row")

    raise RuntimeError(f"No answer after {max_turns} turns")
# Layer 3 is everything else: a bug in this file, a Ctrl+C. Nothing catches it here, on purpose.

run_agent("Open the temple door.")
```

## Points de vigilance

- **Écrivez les erreurs pour le modèle.** Dites ce qui n'allait pas et quoi faire à la place : « la serrure est bloquée par la rouille. Huile-la d'abord : use oil_can avec lock ». Avec « Error 400 », vous obtiendrez le même appel une fois de plus.
- **Ne laissez pas fuiter ce que le modèle ne devrait pas voir.** Les stack traces, les chemins de fichiers et les chaînes de connexion vont dans vos logs. Le modèle reçoit une phrase claire.
- **Les erreurs aussi peuvent tourner en boucle.** Un modèle peut retenter le même mauvais coup jusqu'à épuisement de `maxTurns`. Arrêtez-vous après quelques erreurs d'affilée, ou interceptez-le avec un hook (niveau suivant).
- **Réessayez avec une limite et une attente aléatoire.** Quelques tentatives, un plafond qui double à chaque fois, et un peu d'aléatoire (jitter) pour que cent clients ne reviennent pas tous à la même seconde.
- **Prudence en réessayant des outils qui modifient des choses.** Si un appel qui déplace de l'argent ou réserve une chambre a expiré, il a peut-être abouti quand même. Vérifiez avant de le refaire.

## Patterns liés

- [3 · Concevoir un outil](https://harnesspatterns.dev/fr/patterns/designing-a-tool.md)
- [4 · Quand s'arrêter](https://harnesspatterns.dev/fr/patterns/when-to-stop.md)
- [6 · Les hooks](https://harnesspatterns.dev/fr/patterns/hooks.md)
- [11 · Observabilité et évaluations](https://harnesspatterns.dev/fr/patterns/observability.md)
