第 5 关
循环中的错误
- user
- assistant
- tool_result
- tool_result(错误)
EventBus
问题
一个智能体在玩冒险游戏。它的工具是屏幕底部的动词 look_at 和 use,任务是“打开神殿的门”。模型做了最显而易见的事:用钥匙开门。可是锁生锈了,工具抛出了异常。
面对这种时刻,老式冒险游戏分成两派。有的游戏里,走错一步你就死了:game over,回到上一个存档。另一些游戏里你死不了;游戏只会告诉你为什么行不通,然后你换个办法再试。智能体循环也必须选一派。
如果循环不接住这个错误,Krash 就赢了:异常沿着循环一路往上冒,run() 抛出异常。整次运行就这么丢了,而错误里明明写着该怎么做:“先上油”。唯一能用上这条信息的读者,偏偏没收到。接住它,作为工具结果送回去,模型读到后就会再试一次。这只花一轮。不接住,代价是整次运行。
三种失败
-
模型申请错了
输入无效、JSON 损坏、调用了不存在的工具、钥匙转不动。
交给模型。 错误以标记为错误的工具结果返回。模型读到后,下一轮换个办法试。
-
模型的服务器出了故障
429(请求太多)、5xx(服务器问题)、连接中断。
交给你的代码,重试。 等一会儿再调用,每次都多等一点。模型永远不会知道:根本没有产生需要它读的轮次。
-
谁也修不好
401(API 密钥错误)、你代码里的 bug、用户取消了。
谁都不管。 让它直接抛出。重试没有用,把它瞒着模型,只会让模型去瞎猜。
诀窍在于别把它们混在一起。重试一个 400,只会把同一个坏掉的请求再发一遍。把 503 拿给模型看,只会浪费一轮在它修不了的东西上。而接住你自己代码里的 bug,只会把它藏起来。
代码
使用 astorlm:工具只管抛出异常,接住的活交给 astorlm。你为模型的服务器打开 retry,再监听 EventBus,就能看到两层机制各自在工作。
从零手写:第 2 关的循环,每一层对应一个函数:callModel 带退避地重试服务器,runTool 把每次工具失败都变成模型读得懂的文本,其余的一律放行抛出。在此之上,还有一个小计数器:当模型一直撞同一堵墙时,就放弃。
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const room = { lockOiled: false, doorOpen: false }
const use = tool({
name: 'use',
description: 'Use an inventory item with something in the room.',
schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
execute: async ({ item, target }) => {
if (item === 'oil_can' && target === 'lock') {
room.lockOiled = true
return 'You oil the lock. It looks like it might turn now.'
}
if (item === 'key' && target === 'door') {
// Just throw. astorlm catches it and sends the message back to the model, marked is_error.
if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
room.doorOpen = true
return 'Click! The key turns and the door swings open.'
}
throw new Error(`Nothing happens when you use ${item} with ${target}.`)
},
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [lookAt, use],
// Off by default: without it, a single 429 ends the run.
retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})
// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
if (event.type === 'provider_retry') {
console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
}
})
try {
const last = await agent.run('Open the temple door.')
console.log(last.content)
} catch (err) {
// Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
console.error('The run failed:', err)
}
// Errors in an agent loop, from scratch. Plain fetch, no SDK.
// The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
const room = { lockOiled: false, doorOpen: false }
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { look_at: lookAt, use }
const toolSchemas = [/* one JSON Schema per tool */]
async function use({ item, target }: Record<string, string>): Promise<string> {
if (item === 'oil_can' && target === 'lock') {
room.lockOiled = true
return 'You oil the lock. It looks like it might turn now.'
}
if (item === 'key' && target === 'door') {
// Say what went wrong and what to try instead: this text is all the model will get.
if (!room.lockOiled) throw new Error('the lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
room.doorOpen = true
return 'Click! The key turns and the door swings open.'
}
throw new Error(`nothing happens when you use ${item} with ${target}.`)
}
// Layer 1, the model's server. Retry only what is likely to pass by itself.
async function callModel(messages: Message[], maxAttempts = 3) {
for (let attempt = 1; ; attempt++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
}).catch(() => null) // null: no response at all, the network failed
if (res?.ok) return (await res.json()).choices[0]
// 429 (too many requests), 5xx (server trouble) and network errors tend to pass.
// A 400 or a 401 won't: retrying just sends the same broken request again.
const transient = res === null || res.status === 429 || res.status >= 500
if (!transient || attempt === maxAttempts) throw new Error(`Model call failed: ${res?.status ?? 'network error'}`)
// Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
await new Promise((resolve) => setTimeout(resolve, Math.random() * 1000 * 2 ** (attempt - 1)))
}
}
// Layer 2, the tools. Whatever goes wrong becomes a result the model can read.
async function runTool(call: ToolCall): Promise<{ output: string; isError: boolean }> {
const run = tools[call.function.name]
if (!run) return { output: `Unknown tool: ${call.function.name}. Tools: ${Object.keys(tools).join(', ')}`, isError: true }
try {
const args = JSON.parse(call.function.arguments) // models do send broken JSON now and then
return { output: await run(args), isError: false }
} catch (err) {
return { output: `Error: ${err instanceof Error ? err.message : err}`, isError: true }
}
}
export async function runAgent(prompt: string, maxTurns = 10, maxErrorsInARow = 3): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
let errorsInARow = 0
for (let turn = 1; turn <= maxTurns; turn++) {
const choice = await callModel(messages)
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const { output, isError } = await runTool(call)
// Chat Completions has no is_error field: the text itself has to say it failed.
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
errorsInARow = isError ? errorsInARow + 1 : 0
}
// A model stuck on the same wrong move rarely gets out by itself.
if (errorsInARow >= maxErrorsInARow) throw new Error(`Gave up after ${errorsInARow} tool errors in a row`)
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// Layer 3 is everything else: a bug in this file, an abort. Nothing catches it here, on purpose.
await runAgent('Open the temple door.')
# Errors in an agent loop, from scratch. Standard library only, no SDK.
# The agent plays an adventure game: its tools are the verbs LOOK AT and USE.
import json
import random
import time
import urllib.error
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
room = {"lock_oiled": False, "door_open": False}
def use(item, target):
if item == "oil_can" and target == "lock":
room["lock_oiled"] = True
return "You oil the lock. It looks like it might turn now."
if item == "key" and target == "door":
# Say what went wrong and what to try instead: this text is all the model will get.
if not room["lock_oiled"]:
raise ValueError("the lock is rusted shut and the key won't turn. Oil it first: use oil_can with lock.")
room["door_open"] = True
return "Click! The key turns and the door swings open."
raise ValueError(f"nothing happens when you use {item} with {target}.")
TOOLS = {"look_at": look_at, "use": use}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
def call_model(messages, max_attempts=3):
"""Layer 1, the model's server. Retry only what is likely to pass by itself."""
for attempt in range(1, max_attempts + 1):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
try:
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
except urllib.error.HTTPError as err:
# 429 (too many requests) and 5xx (server trouble) tend to pass.
# A 400 or a 401 won't: retrying just sends the same broken request again.
if (err.code != 429 and err.code < 500) or attempt == max_attempts:
raise
except urllib.error.URLError:
# No response at all: the network failed.
if attempt == max_attempts:
raise
# Exponential backoff with jitter: a random wait under a ceiling that doubles each time.
time.sleep(random.random() * 2 ** (attempt - 1))
def run_tool(call):
"""Layer 2, the tools. Whatever goes wrong becomes a result the model can read."""
name = call["function"]["name"]
run = TOOLS.get(name)
if run is None:
return f"Unknown tool: {name}. Tools: {', '.join(TOOLS)}", True
try:
args = json.loads(call["function"]["arguments"]) # models do send broken JSON now and then
return run(**args), False
except Exception as err:
return f"Error: {err}", True
def run_agent(prompt, max_turns=10, max_errors_in_a_row=3):
messages = [{"role": "user", "content": prompt}]
errors_in_a_row = 0
for _ in range(max_turns):
choice = call_model(messages)
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
output, is_error = run_tool(call)
# Chat Completions has no is_error field: the text itself has to say it failed.
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
errors_in_a_row = errors_in_a_row + 1 if is_error else 0
# A model stuck on the same wrong move rarely gets out by itself.
if errors_in_a_row >= max_errors_in_a_row:
raise RuntimeError(f"Gave up after {errors_in_a_row} tool errors in a row")
raise RuntimeError(f"No answer after {max_turns} turns")
# Layer 3 is everything else: a bug in this file, a Ctrl+C. Nothing catches it here, on purpose.
run_agent("Open the temple door.")
注意事项
- 为模型写错误信息。说清楚哪里错了、应该怎么做:“锁锈死了。先上油:对 lock 使用 oil_can”。只写“Error 400”,你会再收到一次同样的调用。
- 别泄露模型不该看到的东西。堆栈跟踪、文件路径和连接字符串都应该进你的日志。模型只需要一句清楚的话。
-
错误也会陷入循环。模型可能一直重复同一个错误的动作,直到
maxTurns用完。连续出错几次后就停下,或者用钩子拦住它(下一关)。 - 重试要有上限,并加上随机等待。少量几次尝试,每次翻倍的上限,再加一点随机性(jitter),免得一百个客户端在同一秒一起回来。
- 重试会改变状态的工具时要小心。如果一次转账或订房的调用超时了,它可能其实已经成功了。再做一次之前,先确认一下。