跳到正文
astorlm
语言: 简体中文
← 地图

第 3 关

设计一个工具

模型永远看不到你工具的代码。它看到的是一个名称、一段描述和一个 schema,单凭这些就决定要不要用这个工具、要传什么进去。对模型来说,这就是工具的全部。
1/16 手风琴褶数:
  • user
  • assistant
  • tool_result
一只野生生物挡住了去路。它的属性还是 ???。神谕者有四个工具,而它对这些工具所能知道的一切,就是这份招式表:名称、参数和描述。

EventBus

问题

你写一个函数的时候,知道它是干什么的。模型不知道:它拿到的只有工坊门口的那块招牌。招牌写得含糊,模型就只能猜。猜错一次,就要多花一轮、多一个错误,历史记录里也多出一堆 token。更糟的是,一个含糊的工具可能带着一个毫无用处的结果“成功”返回,而谁都没察觉。

在这场战斗里,一只野生生物挡住了路,训练师问该派哪个伙伴出场。神谕者有四个工具,招式菜单展示的正是模型能拿到的每个工具的信息:名称、参数和描述。这就足以让它跳过 do_stuff,因为描述写着“战斗中不可用”而跳过 heal_party,知道必须先查出这只生物的属性,并且第一次就用正确的格式把输入传给 check_matchup。

一个好工具的构成

  • 名称

    不要用 do_stuff,而要用 identify_wild, check_matchup

    具体的动词加名词,能让模型在读其他任何内容之前就知道这个工具是干什么的。

  • 描述

    不要用 "Does stuff.",而要用 它做什么、什么时候用、返回什么。

    这是模型能拿到的唯一文档。要写给一个看不到你代码的人。

  • 参数

    不要用 x: string,而要用 attack 和 defend,各自带描述,并写明格式:一个小写的属性名,比如 water。

    标明哪些是必填,并禁止多余字段,这样错误的输入会尽早失败,而不是做出奇怪的事。

  • 输出

    不要用 整张属性相克表,18 × 18,而要用 一行:grass vs water: 2x, super effective.

    工具返回的一切都会进入历史记录,而模型每一轮都要再读一遍。

  • 错误

    不要用 "Invalid input",而要用 "defend must be one lowercase type, like water. Call identify_wild to get it."

    模型能看懂的错误,就是它能在下一轮修正的错误。

还有一条规则:工具要少,而且要精。你每加一个工具,模型每一轮就要多读一块招牌,也就多了一种选错的可能。

代码

使用 astorlm:tool() 接收一个 Zod schema,把它转换成模型读取的 JSON Schema,并在你的代码运行之前用它校验模型的输入。如果输入不匹配,模型会收到一条清晰的错误,而不是让你的函数崩溃。

从零手写:一个工具就是一份模型读取的 JSON Schema,加上一个模型永远看不到的函数。下面把 Crooky 的 do_stuff 和神谕者在战斗中用过的两个工具并排放在一起。只有 schema 会发送给模型,放在第 2 关每次请求的 tools 字段里。

import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const type = z
  .string()
  .regex(/^[a-z]+$/, 'a type is one lowercase word, like water; call identify_wild to get the wild one')

// Describe the input once with Zod: astorlm turns it into the JSON Schema the model
// reads, and checks the model's input against it before running your code.
const checkMatchup = tool({
  name: 'check_matchup',
  description:
    'How hard one type hits another. Use it to choose which partner to send into a battle. ' +
    'Returns one line: the multiplier (2x, 1x or 0.5x) and what it means.',
  schema: z.object({
    attack: type.describe('The attacking type, one lowercase word, e.g. "grass".'),
    defend: type.describe('The defending type, one lowercase word, e.g. "water".'),
  }),
  execute: async ({ attack, defend }) => {
    const times = typeChart[attack]?.[defend] ?? 1 // your code; the model never sees it
    return `${attack} vs ${defend}: ${times}x${times > 1 ? ', super effective' : ''}`
  },
})

const identifyWild = tool({
  name: 'identify_wild',
  description: 'Identify the wild creature you are facing. Returns its name and type.',
  schema: z.object({}),
  execute: async () => {
    const wild = battle.opponent() // your code
    return `${wild.name} · type: ${wild.type}`
  },
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [checkMatchup, identifyWild],
})

await agent.run('A wild creature appeared! Should I send Emberpup (fire) or Sproutle (grass)?')

像模型那样检查你的工具

  • 只读 schema。把代码藏起来,问问自己:我知道什么时候该调用它、该传什么吗?
  • 盯住最初的几次调用。如果模型总是传错输入,问题几乎总是出在描述上,而不是提示词上。
  • 衡量输出。明明一行就够,工具却返回整张属性相克表,历史记录很快就会被塞满。