第 8 关
按需加载的技能
- user
- assistant
- tool_result
EventBus
问题
一家烘焙店的订单助手需要知道店里的规矩。多大的蛋糕够 20 人吃。无坚果蛋糕只在周五上午烤。定金是多少。外烩托盘怎么报价。货架上每样商品里有什么。这些模型全都不知道。
显而易见的办法是把这些全写下来,贴进 system prompt。第一天很好用。然后手册越来越多,变成了十本。现在每次请求、每一轮都要背着全部手册,不管顾客要的是婚礼蛋糕,还是只想问营业时间。这些你都要付钱,请求越来越慢,唯一要紧的那条规则也被埋在一堆无关的规则里。
这就是 Tomebloat,臃肿的 system prompt。它在喂养第 7 关的 Gulp:一个本来就塞满的提示词,留给对话本身的空间就更少了。
解决方案
把每本手册一分为二。一两行的描述说明这个技能什么时候适用。正文说明这份工作怎么做。只有描述会作为目录放进 system prompt。当某个请求和其中一条描述对上时,模型就申请它的正文,正文从那以后加入对话。
这就是渐进式披露(progressive disclosure):先给出索引,细节只在需要时才展示。常见的格式是每个技能一个文件夹,里面放一个 SKILL.md 文件。开头的 frontmatter(两行 --- 之间的那块)写着名称和描述。它下面的一切都是正文。
---
name: custom-cake-order
description: Cakes made to order. Use when a customer wants a cake baked for them — size by guests, flavors, allergies, lead time, deposit.
---
# Custom cake orders
1. Size by guests: up to 12 → 18 cm, up to 24 → 24 cm, more → two tiers.
2. Flavors: chocolate, vanilla, dulce de leche. Nothing else.
3. Nut allergy → the nut-free line. It only bakes on Friday mornings.
4. Check the calendar before you promise a date. Never less than 72 hours.
5. Quote the 50% deposit, and ask before you book. Never book on your own.
把技能交给模型有三种常见方式。astorlm 通过 skillMode 三种都支持:
-
all从第一次请求开始,每个技能的完整正文都放进 system prompt。
没有额外调用,模型也不可能跳过哪本手册。适合两三个几乎每次请求都要用的简短技能。超过这个数量,就是 Tomebloat 了。
-
on-demandsystem prompt 列出每个技能的名称和描述。模型申请时,load_skill 工具返回完整正文。
能扩展到几十个技能。每加载一个技能要花一次工具调用,而且小模型有时会没加载需要的技能就直接作答。
-
filesystem目录里还给出每个 SKILL.md 的路径,模型用它普通的读文件工具打开。
同样的思路,但不需要专门的工具。Claude Code、Codex 和 Gemini CLI 都是这么做的,所以同一个技能文件夹在它们之间通用。前提是智能体能读文件。
技能不是工具。工具是循环替模型去执行的东西。技能是模型去读的东西,它可以告诉模型该调用哪些工具、按什么顺序,比如那本食谱上写着“承诺日期之前先查日历”。
角色
还是那群熟悉的角色,这次换到了烘焙店的后厨。
- 神谕者 模型
- 出餐台后面的主厨。每一轮它都读面前的一切,自己却什么也不做。
- 单子 目录
- 每个技能一张,挂在出餐台上方的单轨上:一个名称,加一句什么时候用它。它们是 system prompt 的一部分,所以每次请求都带着。
- 架子 技能正文
- 每个技能一本合着的食谱。书越厚,占的 token 越多。在有人去取之前,这些内容一点都不会到达模型。
- 摊开的书 已加载的技能
load_skill把正文作为工具结果返回(那枚绿色书签)。它和其他结果一样进入手风琴,所以在接下来的整个运行中都摊开在出餐台上。- 烤箱 check_calendar
- 一个普通工具。告诉模型去用它的,是那本食谱。
- 条形图 请求大小
- 下一次请求带着多少 token。它的总长度,是把三本书都贴进 system prompt 时会带的量。
在 EventBus 面板里,load_skill 显示为普通的 tool_execution_start 和 tool_execution_end。循环根本不知道技能的存在:对它来说,加载一本手册只是又一次工具调用。
代码
使用 astorlm:把 skillSources 指向一个技能文件夹。默认的 skillMode: 'on-demand' 会替你把目录放进 system prompt,并加上 load_skill 工具。
从零手写:读取每个 SKILL.md,在 system prompt 里给每个技能写一行,再加一个返回正文的 load_skill 工具。第 2 关的循环一点都不用改。
import { OpenAIProvider, createFileSystemSkillSource, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const checkCalendar = tool({
name: 'check_calendar',
description: 'Free baking slots and pickup times for a day, per production line.',
schema: z.object({ day: z.string(), line: z.enum(['regular', 'nut-free']) }),
execute: async ({ day, line }) => bakerySlots(day, line), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [checkCalendar],
// Every folder in ./skills with a SKILL.md is one skill:
// skills/custom-cake-order/SKILL.md, skills/catering-quote/SKILL.md, …
skillSources: [createFileSystemSkillSource({ dir: './skills' })],
// The default. Only each skill's name and description go in the system prompt,
// and the agent gets a load_skill tool to fetch a full body when it needs one.
skillMode: 'on-demand',
maxTurns: 10,
})
agent.on('tool-end', ({ name, output }) => {
if (name === 'load_skill') console.log('loaded:', output.slice(0, 60))
})
const last = await agent.run('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.')
console.log(last.content)
// On-demand skills, from scratch. Plain fetch and node:fs, no SDK.
import { readdirSync, readFileSync } from 'node:fs'
import { join } from 'node:path'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 1. Read the skills: one folder per skill, each with a SKILL.md.
// The frontmatter (the block between the --- lines) holds name and description.
type Skill = { name: string; description: string; body: string }
function readSkill(file: string): Skill {
const [, front = '', body = ''] = readFileSync(file, 'utf8').match(/^---\n([\s\S]*?)\n---\n?([\s\S]*)$/) ?? []
const field = (key: string) => front.match(new RegExp(`^${key}:\\s*(.+)$`, 'm'))?.[1]?.trim() ?? ''
return { name: field('name'), description: field('description'), body: body.trim() }
}
const SKILLS_DIR = './skills'
const skills = new Map(
readdirSync(SKILLS_DIR, { withFileTypes: true })
.filter((entry) => entry.isDirectory())
.map((entry) => readSkill(join(SKILLS_DIR, entry.name, 'SKILL.md')))
.map((skill) => [skill.name, skill]),
)
// 2. The catalog goes in the system prompt: one line per skill, never the body.
const catalog = [...skills.values()].map((s) => `- \`${s.name}\` — ${s.description}`).join('\n')
const SYSTEM = `You are the order assistant of a bakery.
<available-skills>
When the customer's request matches a skill, call load_skill with its name
before you answer, and follow the instructions it returns.
${catalog}
</available-skills>`
// 3. load_skill is a tool like any other. Its result is the body.
const loadSkill = ({ name }: { name: string }): string => {
const skill = skills.get(name)
if (!skill) throw new Error(`Unknown skill "${name}". Available: ${[...skills.keys()].join(', ')}`)
return `<skill name="${skill.name}">\n${skill.body}\n</skill>`
}
type ToolFn = (args: Record<string, string>) => string | Promise<string>
const tools: Record<string, ToolFn> = {
load_skill: (args) => loadSkill({ name: args.name ?? '' }),
check_calendar: checkCalendar, // your code
}
const toolSchemas = [
{
type: 'function',
function: {
name: 'load_skill',
description: "Load a skill's full instructions by name. Skills are listed in <available-skills>.",
parameters: { type: 'object', properties: { name: { type: 'string' } }, required: ['name'] },
},
},
/* check_calendar's schema */
]
// 4. The loop from level 2. Nothing in it knows about skills.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: SYSTEM },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
// A loaded body lands here, in the history, and every later turn resends it.
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
console.log(await runAgent('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.'))
# On-demand skills, from scratch. Standard library only, no SDK.
import json
import re
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
# 1. Read the skills: one folder per skill, each with a SKILL.md.
# The frontmatter (the block between the --- lines) holds name and description.
def read_skill(path):
match = re.match(r"^---\n(.*?)\n---\n?(.*)$", path.read_text(encoding="utf-8"), re.S)
front, body = match.groups() if match else ("", "")
def field(key):
found = re.search(rf"^{key}:\s*(.+)$", front, re.M)
return found.group(1).strip() if found else ""
return {"name": field("name"), "description": field("description"), "body": body.strip()}
SKILLS = {
skill["name"]: skill
for skill in (read_skill(folder / "SKILL.md") for folder in Path("skills").iterdir() if folder.is_dir())
}
# 2. The catalog goes in the system prompt: one line per skill, never the body.
CATALOG = "\n".join(f"- `{s['name']}` — {s['description']}" for s in SKILLS.values())
SYSTEM = f"""You are the order assistant of a bakery.
<available-skills>
When the customer's request matches a skill, call load_skill with its name
before you answer, and follow the instructions it returns.
{CATALOG}
</available-skills>"""
# 3. load_skill is a tool like any other. Its result is the body.
def load_skill(name):
if name not in SKILLS:
raise ValueError(f'Unknown skill "{name}". Available: {", ".join(SKILLS)}')
return f'<skill name="{name}">\n{SKILLS[name]["body"]}\n</skill>'
TOOLS = {"load_skill": load_skill, "check_calendar": check_calendar} # check_calendar: your code
TOOL_SCHEMAS = [
{
"type": "function",
"function": {
"name": "load_skill",
"description": "Load a skill's full instructions by name. Skills are listed in <available-skills>.",
"parameters": {"type": "object", "properties": {"name": {"type": "string"}}, "required": ["name"]},
},
},
# check_calendar's schema
]
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
# 4. The loop from level 2. Nothing in it knows about skills.
def run_agent(prompt, max_turns=10):
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
try:
output = TOOLS[name](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
# A loaded body lands here, in the history, and every later turn resends it.
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
print(run_agent("I need a cake for 20 people this Saturday. Chocolate, and one guest can't have nuts."))
注意事项
-
描述就是触发器。模型只凭描述来挑技能,就像它挑工具一样(第 3 关)。要写什么时候用它,而不是它里面有什么:“当顾客想订做蛋糕时使用”比“蛋糕信息”好得多。在动画里,
allergen-info也提到了过敏;只有描述才能让模型不拿错书。 -
小模型会跳过这一步。有些模型没加载需要的技能就直接作答。在 system prompt 里把话说明白(“作答前先加载对应的技能”),试试 filesystem 模式,或者对每次请求都要用的技能,干脆放在
all里。 - 加载过的技能会一直留着。它的正文在历史记录里,所以之后的每一轮都要为它付费,压缩(第 7 关)之后也可能把它截断。正文要写短,大技能要拆成两个。
-
技能是指令,所以它们就是代码。谁写了
SKILL.md,谁就在操控你的智能体。不要从你无法掌控的文件夹或注册表加载技能(第 14 关)。