Niveau 4
Quand s'arrêter
- user
- assistant
- tool_result
EventBus
Le problème
La boucle du niveau 2 a une sortie normale : le modèle répond sans demander d'outil. Mais un modèle qui ne trouve pas ce qu'il cherche le dit rarement. Il essaie une autre recherche, puis une autre orthographe, puis une autre encore. Chaque tour renvoie tout l'historique, donc chaque tour coûte plus que le précédent.
C'est Loopboros, la boucle qui ne finit jamais. maxTurns l'arrête. Mais regardez à nouveau le monde
1-1 : quand l'horloge s'épuise, la boucle n'échoue pas. Elle s'arrête et vous rend son dernier message, et ce
message est une demande de recherche. Ni réponse, ni erreur.
Les trois sorties
-
Le drapeau
end_turnWorld 1-2
Le modèle décide. Il répond sans demander d'outil. C'est la sortie normale, et la plupart des exécutions devraient finir ici.
Renvoie : Un message avec la réponse en texte.
-
Le tuyau de raccourci
stopOnToolNamesWorld 1-3
Votre code décide, quand le modèle appelle un outil que vous avez marqué. La boucle se ferme juste après que cet outil s'est exécuté sans erreur. S'il échoue, l'erreur revient au modèle et la boucle continue.
Renvoie : Un message avec l'appel d'outil. Son entrée est la réponse, déjà vérifiée par le schéma.
-
L'horloge
maxTurnsWorld 1-1
Personne ne décide. Le budget est épuisé. La boucle s'arrête après ce nombre de tours, quoi que le modèle soit en train de faire. C'est un filet de sécurité, pas une réponse.
Renvoie : Le dernier message, quel qu'il soit. Souvent un appel d'outil.
Un outil terminal ne vous fait pas gagner un tour par rapport à end_turn. Ce qu'il vous donne, c'est
la réponse sous forme de données, vérifiées par un schéma, prêtes à être utilisées par votre app : ici, un
marqueur sur la carte du niveau. Sans stopOnToolNames, après l'exécution de mark_star,
la boucle redemanderait au modèle juste pour qu'il dise « terminé ».
Le code
Avec astorlm : maxTurns règle l'horloge et stopOnToolNames marque les
tuyaux de raccourci. run() renvoie toujours le dernier message : lisez-le pour savoir par quelle
sortie l'exécution s'est terminée.
À partir de zéro : La boucle du niveau 2 avec les trois sorties signalées. Au lieu d'une simple chaîne, elle renvoie comment l'exécution s'est terminée, pour que l'appelant ne confonde pas un timeout avec une réponse.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const searchBlocks = tool({
name: 'search_blocks',
description: 'Search the ? blocks in one area of the level. Returns what each one hides.',
schema: z.object({ area: z.string() }),
execute: async ({ area }) => level.search(area), // your code
})
// The answer as data. Its schema is checked before the loop is allowed to stop.
const markStar = tool({
name: 'mark_star',
description: 'Put a marker on the level map where the star is. Call it once, at the end.',
schema: z.object({ item: z.literal('star'), block: z.number().int().positive() }),
execute: async ({ block }) => map.addMarker(block), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [searchBlocks, markStar],
appendSystemPrompt:
'When you know where the star is, deliver it with mark_star. ' +
"If a few searches come back empty, say you couldn't find it.",
maxTurns: 10, // the clock. Without it, the default is 25.
stopOnToolNames: ['mark_star'], // the warp pipe
})
const last = await agent.run('Is there a hidden star in this level?')
// The loop hands back its last message whichever way it ended. Read it to find out which.
const call = last.content.find((block) => block.type === 'tool_use')
if (!call) {
console.log('end_turn:', last.content) // the flag: a plain text answer
} else if (call.name === 'mark_star') {
console.log('answer:', call.input) // the warp pipe: { item, block }, schema-checked
} else {
// The clock: maxTurns ran out while the model was still asking for tools.
throw new Error(`No answer: the run stopped while asking for ${call.name}`)
}
// The three exits of an agent loop, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, unknown>) => Promise<string>
const tools: Record<string, ToolFn> = { search_blocks: searchBlocks, mark_star: markStar }
const toolSchemas = [/* one JSON Schema per tool */]
// Terminal tools: once one runs without error, the loop is over.
const stopOn = new Set(['mark_star'])
// Say how the run ended, so the caller can tell an answer from a timeout.
type Outcome =
| { stop: 'end_turn'; text: string }
| { stop: 'terminal_tool'; input: Record<string, unknown> }
| { stop: 'max_turns'; turns: number }
export async function runAgent(prompt: string, maxTurns = 10): Promise<Outcome> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
// Exit 1, the flag: no tool calls, the model is done.
if (choice.finish_reason !== 'tool_calls') return { stop: 'end_turn', text: reply.content ?? '' }
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let input: Record<string, unknown> = {}
let output = `Unknown tool: ${call.function.name}`
let failed = true
if (run) {
try {
input = JSON.parse(call.function.arguments)
output = await run(input)
failed = false
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
// Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
// If it failed, the error goes back to the model and the loop keeps going.
if (stopOn.has(call.function.name) && !failed) return { stop: 'terminal_tool', input }
}
}
// Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return { stop: 'max_turns', turns: maxTurns }
}
# The three exits of an agent loop, from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"search_blocks": search_blocks, "mark_star": mark_star}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
# Terminal tools: once one runs without error, the loop is over.
STOP_ON = {"mark_star"}
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, max_turns=10):
"""Returns how the run ended, so the caller can tell an answer from a timeout."""
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
# Exit 1, the flag: no tool calls, the model is done.
if choice["finish_reason"] != "tool_calls":
return {"stop": "end_turn", "text": reply.get("content") or ""}
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])
run = TOOLS.get(name)
failed = run is None
try:
output = run(**args) if run else f"Unknown tool: {name}"
except Exception as err:
failed = True
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
# Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
# If it failed, the error goes back to the model and the loop keeps going.
if name in STOP_ON and not failed:
return {"stop": "terminal_tool", "input": args}
# Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return {"stop": "max_turns", "turns": max_turns}
Points de vigilance
-
Définissez toujours
maxTurns, et dimensionnez-le selon la tâche. Une recherche demande une poignée de tours ; un refactoring peut en demander des dizaines. La valeur par défaut dans astorlm est 25. - Traitez l'horloge comme un échec. Si l'exécution s'est terminée par un appel d'outil, dites à l'utilisateur que vous n'avez pas pu finir, ou réessayez avec un prompt plus clair. Ne lui montrez pas une réponse vide.
- Donnez au modèle un moyen d'abandonner. Dites-lui dans le system prompt de répondre « je ne l'ai pas trouvé » quand une recherche revient vide encore et encore. Un modèle qui a le droit de s'arrêter s'arrête plus tôt.
-
Des tours, ce n'est pas du temps.
maxTurnsne sert à rien face à un seul tour qui ne finit jamais. Pour cela, il vous faut un timeout ou unAbortSignal.