> Niveau 13 de Agent Harness Patterns, un parcours de patterns sur le fonctionnement des agents d'IA. Version web : https://harnesspatterns.dev/fr/patterns/human-in-the-loop · Tous les patterns (en anglais) : https://harnesspatterns.dev/llms.txt

# L'humain dans la boucle

L'agent fait le travail ; une personne garde le dernier mot. Elle approuve ce qui ne peut pas être défait, le corrige en cours de route, et répond quand il hésite.

## Le problème

Un agent qui tourne sans personne est rapide, jusqu'au moment où il fait quelque chose que personne ne voulait : payer la mauvaise facture, envoyer un e-mail à tous les clients, dépenser la dernière pièce pour le mauvais lot. Et un agent qui s'arrête pour tout demander n'est pas plus rapide que de le faire vous-même.

C'est Autopilot Rex : il ne demande jamais, n'attend jamais, ne prend jamais d'appel. Il s'accroche à son premier plan, ignore ce qui a changé depuis et devine là où il aurait dû demander.

## La solution

Placez une personne aux quelques endroits où son jugement compte, et laissez l'agent tourner partout ailleurs. Il y a trois de ces endroits, et ils diffèrent par qui rompt le silence :

- **Approuver**
   L'agent attend
   Avant qu'un outil irréversible ne s'exécute, un hook montre à la personne exactement ce qu'il va faire et attend un oui ou un non.
   Paiements, suppressions, messages aux clients, la dernière pièce : tout ce qu'on ne peut pas reprendre.
- **Réorienter**
   La personne interrompt
   La personne met une correction en file quand elle veut. Au prochain appel d'outil, la boucle annule cet appel et donne la correction au modèle à la place.
   Quelqu'un regarde et change d'avis, ou voit l'agent partir dans la mauvaise direction.
- **Faire remonter**
   L'agent demande
   Un outil comme ask_human qui ne rend pas la main tant qu'une personne n'a pas répondu. Le modèle décide quand l'utiliser.
   Une information manque, deux options se valent, peu de confiance. Le system prompt dit quand demander.

Approuver et réorienter vivent tous deux dans le hook `beforeToolExecution` du niveau 6 : il s'exécute avant chaque outil, et il peut attendre. Pour une approbation, il attend la réponse de la personne ; pour une réorientation, il vérifie si une correction est en file. Dans les deux cas, ce qu'a dit la personne revient au modèle comme résultat de l'outil, donc l'exécution continue avec la nouvelle information au lieu de planter.

## Les personnages

Les mêmes personnages que d'habitude, cette fois dans une salle d'arcade.

- **La voyante** (le modèle): L'Oracle, dans sa cabine. Il lit l'historique et décide de l'appel suivant. Il ne touche jamais la machine.
- **Les commandes** (les outils): `move_claw` est gratuit et peut être annulé. `drop_claw` dépense la dernière pièce : lui, non.
- **Tina** (la personne): Sa bulle dit qui a parlé en premier : **!** elle a interrompu, **?** on lui a demandé, **YES!** elle a approuvé.
- **L'écran** (l'approbation): DROP? YES NO : la pince reste immobile pendant que le hook attend. Rien ne s'exécute tant qu'elle n'a pas décidé.

Dans le panneau EventBus, la réorientation apparaît comme un événement `user_steering`, suivi du résultat de l'appel annulé. Les approbations et les réponses n'ont pas d'événement à elles : ce sont un hook et un outil qui ont pris leur temps.

## Le code

**Avec astorlm :** Un hook d'approbation pour les outils irréversibles, enveloppé par `createSteeringController`, qui ajoute la file de corrections : appelez `steer(text)` depuis votre interface, et elle arrive au prochain appel d'outil. Faire remonter, c'est un outil ordinaire qui attend votre interface.

**À partir de zéro :** Trois vérifications dans l'étape outils de la boucle du niveau 2 : une correction en file, une approbation pour les outils risqués, et un outil qui attend une personne.

**Avec astorlm**

```ts
import { OpenAIProvider, createLocalAgent, createSteeringController, tool } from 'astorlm'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

// Your UI: each of these resolves when the person clicks or types.
declare function approve(what: string): Promise<boolean> // shows YES / NO
declare function answer(question: string): Promise<string> // shows a text box

// 1. ESCALATE: a tool the model calls when it isn't sure. It waits for a person.
const askKid = tool({
  name: 'ask_kid',
  description: 'Ask Tina when you are not sure what she wants. Waits for her answer.',
  schema: z.object({ question: z.string() }),
  execute: async ({ question }) => answer(question),
})

// 2. APPROVE: irreversible tools wait for a yes before they run.
const NEEDS_APPROVAL = new Set(['drop_claw'])
const steering = createSteeringController({
  beforeToolExecution: async ({ toolName, input }) => {
    if (!NEEDS_APPROVAL.has(toolName)) return { authorize: true }
    const ok = await approve(`${toolName}(${JSON.stringify(input)})`) // show the real call
    return ok ? { authorize: true } : { authorize: false, mockResult: 'Tina said no. Ask her what to do.' }
  },
})

const agent = await createLocalAgent({
  provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
  systemPrompt: 'You work a claw machine for Tina. If you are not sure which prize she means, ask her before acting.',
  tools: [moveClaw, dropClaw, askKid], // moveClaw, dropClaw: your code
  hooks: steering.hooks, // the steering controller wraps the approval hook
  maxTurns: 20,
})

// 3. STEER: the person can correct the agent at any moment, e.g. from a button.
// The loop cancels the next tool call and hands the model this feedback instead.
onTinaShouts((text) => steering.steer(text)) // your UI: 'No, wait! The penguin!'

const result = await agent.run('Get me the bear!')
console.log(result.content)
```

**TypeScript**

```ts
// Human in the loop, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

// Your UI: each of these resolves when the person clicks or types.
declare function approve(what: string): Promise<boolean>
declare function answer(question: string): Promise<string>

type ToolFn = (args: Record<string, string | number>) => Promise<string>
const tools: Record<string, ToolFn> = {
  move_claw: moveClaw, // your code
  drop_claw: dropClaw, // your code
  // 1. ESCALATE: the model asks, a person answers.
  ask_kid: ({ question }) => answer(String(question)),
}
const toolSchemas = [/* one JSON Schema per tool: move_claw(to), drop_claw(), ask_kid(question) */]
const NEEDS_APPROVAL = new Set(['drop_claw'])

// 3. STEER: a person can queue a correction at any time, e.g. from a button.
let steer: string | null = null
export const queueCorrection = (text: string) => {
  steer = text
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

export async function runAgent(prompt: string, maxTurns = 20): Promise<string> {
  const messages: Message[] = [
    { role: 'system', content: 'You work a claw machine for Tina. If you are not sure which prize she means, ask her before acting.' },
    { role: 'user', content: prompt },
  ]
  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const name = call.function.name
      const args = JSON.parse(call.function.arguments)
      let output: string
      if (steer !== null) {
        // The tool boundary: a queued correction cancels this call, and the model reads it instead.
        output = `Cancelled. The person says: ${steer}`
        steer = null
      } else if (NEEDS_APPROVAL.has(name) && !(await approve(`${name}(${call.function.arguments})`))) {
        // 2. APPROVE: irreversible tools wait here for a yes.
        output = 'Tina said no. Ask her what to do.'
      } else {
        output = tools[name] ? await tools[name](args) : `Unknown tool: ${name}`
      }
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

console.log(await runAgent('Get me the bear!'))
```

**Python**

```python
# Human in the loop, from scratch. Standard library only, no SDK.
import json
import queue
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

def post(path, payload):
    request = urllib.request.Request(
        f"{LLM['base_url']}{path}",
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

# In a terminal the person is input(); in an app, whatever your UI sends back.
def approve(what):
    return input(f"Approve {what}? [y/N] ").strip().lower() == "y"

def ask_kid(question):  # 1. ESCALATE: the model asks, a person answers.
    return input(f"The agent asks: {question}\n> ")

TOOLS = {"move_claw": move_claw, "drop_claw": drop_claw, "ask_kid": ask_kid}  # move_claw, drop_claw: your code
TOOL_SCHEMAS = [...]  # one JSON Schema per tool: move_claw(to), drop_claw(), ask_kid(question)
NEEDS_APPROVAL = {"drop_claw"}

# 3. STEER: another thread (a UI, a chat) can queue a correction at any time.
CORRECTIONS = queue.Queue()

def run_agent(prompt, max_turns=20):
    messages = [
        {"role": "system", "content": "You work a claw machine for Tina. If you are not sure which prize she means, ask her before acting."},
        {"role": "user", "content": prompt},
    ]
    for _ in range(max_turns):
        choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            args = json.loads(call["function"]["arguments"])
            if not CORRECTIONS.empty():
                # The tool boundary: a queued correction cancels this call, and the model reads it instead.
                output = f"Cancelled. The person says: {CORRECTIONS.get()}"
            elif name in NEEDS_APPROVAL and not approve(f"{name}({args})"):
                # 2. APPROVE: irreversible tools wait here for a yes.
                output = "Tina said no. Ask her what to do."
            else:
                output = TOOLS[name](**args)
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

print(run_agent("Get me the bear!"))
```

## Points de vigilance

- **Demandez moins, et ça compte plus.** N'approuvez que l'irréversible ou le coûteux. Une personne qui clique sur YES vingt fois par heure arrête de lire, et l'approbation devient une formalité.
- **Montrez le vrai appel.** « Descendre ? » ne suffit pas. Montrez l'outil et ses arguments exacts, le montant, le destinataire, le lot, pour que la personne approuve ce qui va réellement s'exécuter.
- **Décidez de ce qui se passe quand personne ne répond.** Une personne peut s'éloigner. Fixez un timeout, et choisissez la valeur par défaut sûre : en général refuser, et dire au modèle pourquoi.
- **Les corrections arrivent au prochain appel d'outil.** La réorientation ne peut pas arrêter un outil en pleine exécution ni une réponse au milieu d'un mot. Si l'agent ne fait qu'écrire du texte, la correction attend. Pour un arrêt net, interrompez l'exécution.
- **Gardez une trace.** Journalisez qui a approuvé quoi, quand, et ce qu'il a vu. Quand quelque chose tourne mal, « c'est l'agent qui l'a fait » n'est jamais toute l'histoire.

## Patterns liés

- [6 · Les hooks](https://harnesspatterns.dev/fr/patterns/hooks.md)
- [12 · Planifier et réfléchir](https://harnesspatterns.dev/fr/patterns/plan-and-reflect.md)
- [5 · Les erreurs dans la boucle](https://harnesspatterns.dev/fr/patterns/errors-in-the-loop.md)
- [14 · Sécurité et sandboxing](https://harnesspatterns.dev/fr/patterns/security.md)
- [16 · Les agents proactifs](https://harnesspatterns.dev/fr/patterns/proactive-agents.md)
