> Agent Harness Patterns 第 8 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/skills · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 按需加载的技能

技能（skill）是某一类工作的手册。智能体始终能看到手册的清单，只有在请求需要时才去读其中一本。

## 问题

一家烘焙店的订单助手需要知道店里的规矩。多大的蛋糕够 20 人吃。无坚果蛋糕只在周五上午烤。定金是多少。外烩托盘怎么报价。货架上每样商品里有什么。这些模型全都不知道。

显而易见的办法是把这些全写下来，贴进 system prompt。第一天很好用。然后手册越来越多，变成了十本。现在每次请求、每一轮都要背着全部手册，不管顾客要的是婚礼蛋糕，还是只想问营业时间。这些你都要付钱，请求越来越慢，唯一要紧的那条规则也被埋在一堆无关的规则里。

这就是 Tomebloat，臃肿的 system prompt。它在喂养第 7 关的 Gulp：一个本来就塞满的提示词，留给对话本身的空间就更少了。

## 解决方案

把每本手册一分为二。一两行的**描述**说明这个技能什么时候适用。**正文**说明这份工作怎么做。只有描述会作为目录放进 system prompt。当某个请求和其中一条描述对上时，模型就申请它的正文，正文从那以后加入对话。

这就是渐进式披露（progressive disclosure）：先给出索引，细节只在需要时才展示。常见的格式是每个技能一个文件夹，里面放一个 `SKILL.md` 文件。开头的 frontmatter（两行 `---` 之间的那块）写着名称和描述。它下面的一切都是正文。

**skills/custom-cake-order/SKILL.md**

```md
---
name: custom-cake-order
description: Cakes made to order. Use when a customer wants a cake baked for them — size by guests, flavors, allergies, lead time, deposit.
---

# Custom cake orders

1. Size by guests: up to 12 → 18 cm, up to 24 → 24 cm, more → two tiers.
2. Flavors: chocolate, vanilla, dulce de leche. Nothing else.
3. Nut allergy → the nut-free line. It only bakes on Friday mornings.
4. Check the calendar before you promise a date. Never less than 72 hours.
5. Quote the 50% deposit, and ask before you book. Never book on your own.
```

把技能交给模型有三种常见方式。astorlm 通过 `skillMode` 三种都支持：

- `all`
   从第一次请求开始，每个技能的完整正文都放进 system prompt。
   没有额外调用，模型也不可能跳过哪本手册。适合两三个几乎每次请求都要用的简短技能。超过这个数量，就是 Tomebloat 了。
- `on-demand`
   system prompt 列出每个技能的名称和描述。模型申请时，load_skill 工具返回完整正文。
   能扩展到几十个技能。每加载一个技能要花一次工具调用，而且小模型有时会没加载需要的技能就直接作答。
- `filesystem`
   目录里还给出每个 SKILL.md 的路径，模型用它普通的读文件工具打开。
   同样的思路，但不需要专门的工具。Claude Code、Codex 和 Gemini CLI 都是这么做的，所以同一个技能文件夹在它们之间通用。前提是智能体能读文件。

技能不是工具。工具是循环替模型去执行的东西。技能是模型去读的东西，它可以告诉模型该调用哪些工具、按什么顺序，比如那本食谱上写着“承诺日期之前先查日历”。

## 角色

还是那群熟悉的角色，这次换到了烘焙店的后厨。

- **神谕者** (模型): 出餐台后面的主厨。每一轮它都读面前的一切，自己却什么也不做。
- **单子** (目录): 每个技能一张，挂在出餐台上方的单轨上：一个名称，加一句什么时候用它。它们是 system prompt 的一部分，所以每次请求都带着。
- **架子** (技能正文): 每个技能一本合着的食谱。书越厚，占的 token 越多。在有人去取之前，这些内容一点都不会到达模型。
- **摊开的书** (已加载的技能): `load_skill` 把正文作为工具结果返回（那枚绿色书签）。它和其他结果一样进入手风琴，所以在接下来的整个运行中都摊开在出餐台上。
- **烤箱** (check_calendar): 一个普通工具。告诉模型去用它的，是那本食谱。
- **条形图** (请求大小): 下一次请求带着多少 token。它的总长度，是把三本书都贴进 system prompt 时会带的量。

在 EventBus 面板里，`load_skill` 显示为普通的 `tool_execution_start` 和 `tool_execution_end`。循环根本不知道技能的存在：对它来说，加载一本手册只是又一次工具调用。

## 代码

**使用 astorlm：**把 `skillSources` 指向一个技能文件夹。默认的 `skillMode: 'on-demand'` 会替你把目录放进 system prompt，并加上 `load_skill` 工具。

**从零手写：**读取每个 `SKILL.md`，在 system prompt 里给每个技能写一行，再加一个返回正文的 `load_skill` 工具。第 2 关的循环一点都不用改。

**使用 astorlm**

```ts
import { OpenAIProvider, createFileSystemSkillSource, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const checkCalendar = tool({
  name: 'check_calendar',
  description: 'Free baking slots and pickup times for a day, per production line.',
  schema: z.object({ day: z.string(), line: z.enum(['regular', 'nut-free']) }),
  execute: async ({ day, line }) => bakerySlots(day, line), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [checkCalendar],
  // Every folder in ./skills with a SKILL.md is one skill:
  //   skills/custom-cake-order/SKILL.md, skills/catering-quote/SKILL.md, …
  skillSources: [createFileSystemSkillSource({ dir: './skills' })],
  // The default. Only each skill's name and description go in the system prompt,
  // and the agent gets a load_skill tool to fetch a full body when it needs one.
  skillMode: 'on-demand',
  maxTurns: 10,
})

agent.on('tool-end', ({ name, output }) => {
  if (name === 'load_skill') console.log('loaded:', output.slice(0, 60))
})

const last = await agent.run('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.')
console.log(last.content)
```

**TypeScript**

```ts
// On-demand skills, from scratch. Plain fetch and node:fs, no SDK.
import { readdirSync, readFileSync } from 'node:fs'
import { join } from 'node:path'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// 1. Read the skills: one folder per skill, each with a SKILL.md.
//    The frontmatter (the block between the --- lines) holds name and description.
type Skill = { name: string; description: string; body: string }

function readSkill(file: string): Skill {
  const [, front = '', body = ''] = readFileSync(file, 'utf8').match(/^---\n([\s\S]*?)\n---\n?([\s\S]*)$/) ?? []
  const field = (key: string) => front.match(new RegExp(`^${key}:\\s*(.+)$`, 'm'))?.[1]?.trim() ?? ''
  return { name: field('name'), description: field('description'), body: body.trim() }
}

const SKILLS_DIR = './skills'
const skills = new Map(
  readdirSync(SKILLS_DIR, { withFileTypes: true })
    .filter((entry) => entry.isDirectory())
    .map((entry) => readSkill(join(SKILLS_DIR, entry.name, 'SKILL.md')))
    .map((skill) => [skill.name, skill]),
)

// 2. The catalog goes in the system prompt: one line per skill, never the body.
const catalog = [...skills.values()].map((s) => `- \`${s.name}\` — ${s.description}`).join('\n')
const SYSTEM = `You are the order assistant of a bakery.

<available-skills>
When the customer's request matches a skill, call load_skill with its name
before you answer, and follow the instructions it returns.
${catalog}
</available-skills>`

// 3. load_skill is a tool like any other. Its result is the body.
const loadSkill = ({ name }: { name: string }): string => {
  const skill = skills.get(name)
  if (!skill) throw new Error(`Unknown skill "${name}". Available: ${[...skills.keys()].join(', ')}`)
  return `<skill name="${skill.name}">\n${skill.body}\n</skill>`
}

type ToolFn = (args: Record<string, string>) => string | Promise<string>
const tools: Record<string, ToolFn> = {
  load_skill: (args) => loadSkill({ name: args.name ?? '' }),
  check_calendar: checkCalendar, // your code
}
const toolSchemas = [
  {
    type: 'function',
    function: {
      name: 'load_skill',
      description: "Load a skill's full instructions by name. Skills are listed in <available-skills>.",
      parameters: { type: 'object', properties: { name: { type: 'string' } }, required: ['name'] },
    },
  },
  /* check_calendar's schema */
]

// 4. The loop from level 2. Nothing in it knows about skills.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  const messages: Message[] = [
    { role: 'system', content: SYSTEM },
    { role: 'user', content: prompt },
  ]

  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      try {
        if (run) output = await run(JSON.parse(call.function.arguments))
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      // A loaded body lands here, in the history, and every later turn resends it.
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

console.log(await runAgent('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.'))
```

**Python**

```python
# On-demand skills, from scratch. Standard library only, no SDK.
import json
import re
import urllib.request
from pathlib import Path

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

# 1. Read the skills: one folder per skill, each with a SKILL.md.
#    The frontmatter (the block between the --- lines) holds name and description.
def read_skill(path):
    match = re.match(r"^---\n(.*?)\n---\n?(.*)$", path.read_text(encoding="utf-8"), re.S)
    front, body = match.groups() if match else ("", "")

    def field(key):
        found = re.search(rf"^{key}:\s*(.+)$", front, re.M)
        return found.group(1).strip() if found else ""

    return {"name": field("name"), "description": field("description"), "body": body.strip()}

SKILLS = {
    skill["name"]: skill
    for skill in (read_skill(folder / "SKILL.md") for folder in Path("skills").iterdir() if folder.is_dir())
}

# 2. The catalog goes in the system prompt: one line per skill, never the body.
CATALOG = "\n".join(f"- `{s['name']}` — {s['description']}" for s in SKILLS.values())
SYSTEM = f"""You are the order assistant of a bakery.

<available-skills>
When the customer's request matches a skill, call load_skill with its name
before you answer, and follow the instructions it returns.
{CATALOG}
</available-skills>"""

# 3. load_skill is a tool like any other. Its result is the body.
def load_skill(name):
    if name not in SKILLS:
        raise ValueError(f'Unknown skill "{name}". Available: {", ".join(SKILLS)}')
    return f'<skill name="{name}">\n{SKILLS[name]["body"]}\n</skill>'

TOOLS = {"load_skill": load_skill, "check_calendar": check_calendar}  # check_calendar: your code
TOOL_SCHEMAS = [
    {
        "type": "function",
        "function": {
            "name": "load_skill",
            "description": "Load a skill's full instructions by name. Skills are listed in <available-skills>.",
            "parameters": {"type": "object", "properties": {"name": {"type": "string"}}, "required": ["name"]},
        },
    },
    # check_calendar's schema
]

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

# 4. The loop from level 2. Nothing in it knows about skills.
def run_agent(prompt, max_turns=10):
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]

    for _ in range(max_turns):
        choice = chat(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            try:
                output = TOOLS[name](**json.loads(call["function"]["arguments"]))
            except Exception as err:
                output = f"Error: {err}"
            # A loaded body lands here, in the history, and every later turn resends it.
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

print(run_agent("I need a cake for 20 people this Saturday. Chocolate, and one guest can't have nuts."))
```

## 注意事项

- **描述就是触发器。**模型只凭描述来挑技能，就像它挑工具一样（第 3 关）。要写什么时候用它，而不是它里面有什么：“当顾客想订做蛋糕时使用”比“蛋糕信息”好得多。在动画里，`allergen-info` 也提到了过敏；只有描述才能让模型不拿错书。
- **小模型会跳过这一步。**有些模型没加载需要的技能就直接作答。在 system prompt 里把话说明白（“作答前先加载对应的技能”），试试 filesystem 模式，或者对每次请求都要用的技能，干脆放在 `all` 里。
- **加载过的技能会一直留着。**它的正文在历史记录里，所以之后的每一轮都要为它付费，压缩（第 7 关）之后也可能把它截断。正文要写短，大技能要拆成两个。
- **技能是指令，所以它们就是代码。**谁写了 `SKILL.md`，谁就在操控你的智能体。不要从你无法掌控的文件夹或注册表加载技能（第 14 关）。

## 相关模式

- [0 · 你的工具箱](https://harnesspatterns.dev/zh/patterns/your-toolkit.md)
- [3 · 设计一个工具](https://harnesspatterns.dev/zh/patterns/designing-a-tool.md)
- [7 · 背包装满了](https://harnesspatterns.dev/zh/patterns/compaction.md)
- [9 · 记忆](https://harnesspatterns.dev/zh/patterns/memory.md)
- [14 · 安全与沙箱](https://harnesspatterns.dev/zh/patterns/security.md)
