Nível 4
Quando parar
- user
- assistant
- tool_result
EventBus
O problema
O loop do nível 2 tem uma saída normal: o modelo responde sem pedir uma ferramenta. Mas um modelo que não acha o que procura raramente admite isso. Ele tenta outra busca, depois outra grafia, depois mais outra. Cada volta reenvia o histórico inteiro, então cada volta custa mais que a anterior.
Esse é o Loopboros, o loop que nunca termina. maxTurns o detém. Mas olhe de novo o mundo 1-1: quando
o relógio acaba, o loop não falha. Ele para e devolve a última mensagem, e essa mensagem é um pedido de busca. Nem
resposta, nem erro.
As três saídas
-
A bandeira
end_turnWorld 1-2
O modelo decide. Ele responde sem pedir nenhuma ferramenta. Essa é a saída normal, e a maioria das execuções deveria terminar aqui.
Devolve: Uma mensagem com a resposta em texto.
-
O cano de atalho
stopOnToolNamesWorld 1-3
O seu código decide, quando o modelo chama uma ferramenta que você marcou. O loop se fecha logo depois que essa ferramenta roda sem erro. Se ela falhar, o erro volta para o modelo e o loop continua.
Devolve: Uma mensagem com a chamada da ferramenta. A entrada dela é a resposta, já validada contra o schema.
-
O relógio
maxTurnsWorld 1-1
Ninguém decide. O orçamento acaba. O loop para depois desse número de turnos, seja lá o que o modelo estivesse fazendo. É uma rede de segurança, não uma resposta.
Devolve: O que quer que tenha sido a última mensagem. Muitas vezes, uma chamada de ferramenta.
Uma ferramenta terminal não economiza um turno em relação ao end_turn. O que ela te dá é a resposta
como dado, validada por um schema, pronta para o seu app usar: aqui, um marcador no mapa da fase. Sem
stopOnToolNames, depois que mark_star rodasse, o loop perguntaria ao modelo de novo só
para ele dizer "pronto".
O código
Com astorlm: maxTurns define o relógio e stopOnToolNames marca os
canos de atalho. run() sempre devolve a última mensagem, então leia-a para descobrir por qual saída a
execução terminou.
Do zero: O loop do nível 2 com as três saídas marcadas. Em vez de uma string pura, ele devolve como a execução terminou, para quem chama não confundir um timeout com uma resposta.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const searchBlocks = tool({
name: 'search_blocks',
description: 'Search the ? blocks in one area of the level. Returns what each one hides.',
schema: z.object({ area: z.string() }),
execute: async ({ area }) => level.search(area), // your code
})
// The answer as data. Its schema is checked before the loop is allowed to stop.
const markStar = tool({
name: 'mark_star',
description: 'Put a marker on the level map where the star is. Call it once, at the end.',
schema: z.object({ item: z.literal('star'), block: z.number().int().positive() }),
execute: async ({ block }) => map.addMarker(block), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [searchBlocks, markStar],
appendSystemPrompt:
'When you know where the star is, deliver it with mark_star. ' +
"If a few searches come back empty, say you couldn't find it.",
maxTurns: 10, // the clock. Without it, the default is 25.
stopOnToolNames: ['mark_star'], // the warp pipe
})
const last = await agent.run('Is there a hidden star in this level?')
// The loop hands back its last message whichever way it ended. Read it to find out which.
const call = last.content.find((block) => block.type === 'tool_use')
if (!call) {
console.log('end_turn:', last.content) // the flag: a plain text answer
} else if (call.name === 'mark_star') {
console.log('answer:', call.input) // the warp pipe: { item, block }, schema-checked
} else {
// The clock: maxTurns ran out while the model was still asking for tools.
throw new Error(`No answer: the run stopped while asking for ${call.name}`)
}
// The three exits of an agent loop, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, unknown>) => Promise<string>
const tools: Record<string, ToolFn> = { search_blocks: searchBlocks, mark_star: markStar }
const toolSchemas = [/* one JSON Schema per tool */]
// Terminal tools: once one runs without error, the loop is over.
const stopOn = new Set(['mark_star'])
// Say how the run ended, so the caller can tell an answer from a timeout.
type Outcome =
| { stop: 'end_turn'; text: string }
| { stop: 'terminal_tool'; input: Record<string, unknown> }
| { stop: 'max_turns'; turns: number }
export async function runAgent(prompt: string, maxTurns = 10): Promise<Outcome> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
// Exit 1, the flag: no tool calls, the model is done.
if (choice.finish_reason !== 'tool_calls') return { stop: 'end_turn', text: reply.content ?? '' }
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let input: Record<string, unknown> = {}
let output = `Unknown tool: ${call.function.name}`
let failed = true
if (run) {
try {
input = JSON.parse(call.function.arguments)
output = await run(input)
failed = false
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
// Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
// If it failed, the error goes back to the model and the loop keeps going.
if (stopOn.has(call.function.name) && !failed) return { stop: 'terminal_tool', input }
}
}
// Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return { stop: 'max_turns', turns: maxTurns }
}
# The three exits of an agent loop, from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"search_blocks": search_blocks, "mark_star": mark_star}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
# Terminal tools: once one runs without error, the loop is over.
STOP_ON = {"mark_star"}
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, max_turns=10):
"""Returns how the run ended, so the caller can tell an answer from a timeout."""
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
# Exit 1, the flag: no tool calls, the model is done.
if choice["finish_reason"] != "tool_calls":
return {"stop": "end_turn", "text": reply.get("content") or ""}
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])
run = TOOLS.get(name)
failed = run is None
try:
output = run(**args) if run else f"Unknown tool: {name}"
except Exception as err:
failed = True
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
# Exit 2, the warp pipe: a terminal tool ran fine. Its input is the answer.
# If it failed, the error goes back to the model and the loop keeps going.
if name in STOP_ON and not failed:
return {"stop": "terminal_tool", "input": args}
# Exit 3, the clock: out of turns. Say so out loud; don't pass a tool call off as an answer.
return {"stop": "max_turns", "turns": max_turns}
O que observar
-
Sempre defina
maxTurns, e dimensione-o para a tarefa. Uma consulta precisa de um punhado de turnos; um refactor pode precisar de dezenas. O padrão no astorlm é 25. - Trate o relógio como uma falha. Se a execução terminou com uma chamada de ferramenta, diga ao usuário que não deu para terminar, ou tente de novo com um prompt mais claro. Não mostre uma resposta vazia.
- Dê ao modelo um jeito de desistir. Diga no system prompt para ele responder "não encontrei" quando uma busca continuar voltando vazia. Um modelo que tem permissão para parar, para mais cedo.
-
Turnos não são tempo.
maxTurnsnão ajuda com um único turno que nunca termina. Para isso você precisa de um timeout ou de umAbortSignal.