跳到正文
astorlm
语言: 简体中文
← 地图

第 5 关

循环中的错误

工具会出错。模型会传来错误的输入。服务器会宕机一秒。对每一种失败,智能体循环都要决定由谁来处理:模型、你的代码,还是谁都不管。
1/40 手风琴褶数:
  • user
  • assistant
  • tool_result
  • tool_result(错误)
一个正在玩冒险游戏的智能体:它的工具就是屏幕下方的动词 LOOK AT 和 USE。这个循环是手写的,执行工具时没有 try/catch。

EventBus

问题

一个智能体在玩冒险游戏。它的工具是屏幕底部的动词 look_at 和 use,任务是“打开神殿的门”。模型做了最显而易见的事:用钥匙开门。可是锁生锈了,工具抛出了异常。

面对这种时刻,老式冒险游戏分成两派。有的游戏里,走错一步你就死了:game over,回到上一个存档。另一些游戏里你死不了;游戏只会告诉你为什么行不通,然后你换个办法再试。智能体循环也必须选一派。

如果循环不接住这个错误,Krash 就赢了:异常沿着循环一路往上冒,run() 抛出异常。整次运行就这么丢了,而错误里明明写着该怎么做:“先上油”。唯一能用上这条信息的读者,偏偏没收到。接住它,作为工具结果送回去,模型读到后就会再试一次。这只花一轮。不接住,代价是整次运行。

三种失败

  • 模型申请错了

    输入无效、JSON 损坏、调用了不存在的工具、钥匙转不动。

    交给模型。 错误以标记为错误的工具结果返回。模型读到后,下一轮换个办法试。

  • 模型的服务器出了故障

    429(请求太多)、5xx(服务器问题)、连接中断。

    交给你的代码,重试。 等一会儿再调用,每次都多等一点。模型永远不会知道:根本没有产生需要它读的轮次。

  • 谁也修不好

    401(API 密钥错误)、你代码里的 bug、用户取消了。

    谁都不管。 让它直接抛出。重试没有用,把它瞒着模型,只会让模型去瞎猜。

诀窍在于别把它们混在一起。重试一个 400,只会把同一个坏掉的请求再发一遍。把 503 拿给模型看,只会浪费一轮在它修不了的东西上。而接住你自己代码里的 bug,只会把它藏起来。

代码

使用 astorlm:工具只管抛出异常,接住的活交给 astorlm。你为模型的服务器打开 retry,再监听 EventBus,就能看到两层机制各自在工作。

从零手写:第 2 关的循环,每一层对应一个函数:callModel 带退避地重试服务器,runTool 把每次工具失败都变成模型读得懂的文本,其余的一律放行抛出。在此之上,还有一个小计数器:当模型一直撞同一堵墙时,就放弃。

import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const room = { lockOiled: false, doorOpen: false }

const use = tool({
  name: 'use',
  description: 'Use an inventory item with something in the room.',
  schema: z.object({ item: z.enum(['key', 'oil_can']), target: z.string() }),
  execute: async ({ item, target }) => {
    if (item === 'oil_can' && target === 'lock') {
      room.lockOiled = true
      return 'You oil the lock. It looks like it might turn now.'
    }
    if (item === 'key' && target === 'door') {
      // Just throw. astorlm catches it and sends the message back to the model, marked is_error.
      if (!room.lockOiled) throw new Error('The lock is rusted shut and the key won’t turn. Oil it first: use oil_can with lock.')
      room.doorOpen = true
      return 'Click! The key turns and the door swings open.'
    }
    throw new Error(`Nothing happens when you use ${item} with ${target}.`)
  },
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [lookAt, use],
  // Off by default: without it, a single 429 ends the run.
  retry: { maxAttempts: 3, baseDelayMs: 1000 }, // waits up to 1s, then up to 2s
  maxTurns: 10, // also the ceiling for a model stuck on the same wrong move
})

// Watch both layers as they happen.
agent.on('tool-end', ({ name, output, isError }) => {
  if (isError) console.warn(`${name} failed, and the model will read why: ${output}`)
})
agent.on('event', (event) => {
  if (event.type === 'provider_retry') {
    console.warn(`Model call failed. Attempt ${event.attempt + 1}/${event.maxAttempts} in ${event.delayMs}ms`)
  }
})

try {
  const last = await agent.run('Open the temple door.')
  console.log(last.content)
} catch (err) {
  // Only what nobody could handle gets here: retries used up, a 401, a bug in your code.
  console.error('The run failed:', err)
}

注意事项

  • 为模型写错误信息。说清楚哪里错了、应该怎么做:“锁锈死了。先上油:对 lock 使用 oil_can”。只写“Error 400”,你会再收到一次同样的调用。
  • 别泄露模型不该看到的东西。堆栈跟踪、文件路径和连接字符串都应该进你的日志。模型只需要一句清楚的话。
  • 错误也会陷入循环。模型可能一直重复同一个错误的动作,直到 maxTurns 用完。连续出错几次后就停下,或者用钩子拦住它(下一关)。
  • 重试要有上限,并加上随机等待。少量几次尝试,每次翻倍的上限,再加一点随机性(jitter),免得一百个客户端在同一秒一起回来。
  • 重试会改变状态的工具时要小心。如果一次转账或订房的调用超时了,它可能其实已经成功了。再做一次之前,先确认一下。