> Level 8 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/skills · All patterns: https://harnesspatterns.dev/llms.txt

# On-demand skills

A skill is a manual for one kind of job. The agent always sees the list of manuals, and reads one only when a request calls for it.

## The problem

A bakery’s order assistant needs to know the house rules. Which cake size feeds 20 people. That nut-free cakes only bake on Friday mornings. How much the deposit is. How to quote a catering tray. What goes into each product already on the shelves. The model knows none of it.

The obvious fix is to write it all down and paste it into the system prompt. It works on day one. Then the manuals grow, and there are ten of them. Every request now carries every manual, on every turn, whether the customer wants a wedding cake or just the opening hours. You pay for all of it, the requests get slower, and the one rule that matters is buried among the ones that don’t.

That’s Tomebloat, the bloated system prompt. It feeds Gulp from level 7: a prompt that’s already full leaves less room for the conversation itself.

## The solution

Split each manual in two. A **description** of one or two lines says when the skill applies. The **body** says how to do the job. Only the descriptions go in the system prompt, as a catalog. When a request matches one, the model asks for its body, and the body joins the conversation from then on.

That’s progressive disclosure: show the index first, and the detail only when it’s needed. The usual format is a folder per skill with a `SKILL.md` file. The frontmatter at the top (the block between the `---` lines) holds the name and the description. Everything below it is the body.

**skills/custom-cake-order/SKILL.md**

```md
---
name: custom-cake-order
description: Cakes made to order. Use when a customer wants a cake baked for them — size by guests, flavors, allergies, lead time, deposit.
---

# Custom cake orders

1. Size by guests: up to 12 → 18 cm, up to 24 → 24 cm, more → two tiers.
2. Flavors: chocolate, vanilla, dulce de leche. Nothing else.
3. Nut allergy → the nut-free line. It only bakes on Friday mornings.
4. Check the calendar before you promise a date. Never less than 72 hours.
5. Quote the 50% deposit, and ask before you book. Never book on your own.
```

There are three common ways to hand skills to the model. astorlm supports all three with `skillMode`:

- `all`
   Every skill’s full body goes into the system prompt, from the first request.
   No extra calls, and the model can’t skip a manual. Fine for two or three short skills that almost every request needs. Past that, it’s Tomebloat.
- `on-demand`
   The system prompt lists each skill’s name and description. A load_skill tool returns a full body when the model asks for it.
   Scales to dozens of skills. Costs one tool call per skill loaded, and small models sometimes answer without loading the skill they needed.
- `filesystem`
   The catalog also gives each SKILL.md’s path, and the model opens it with its ordinary file-reading tool.
   The same idea with no special tool. It’s how Claude Code, Codex and Gemini CLI do it, so the same skills folder works in all of them. Needs an agent that can read files.

A skill is not a tool. A tool is something the loop runs for the model. A skill is something the model reads, and it can tell the model which tools to call and in what order, like the recipe that says “check the calendar before you promise a date”.

## The cast

Same cast as always, in a bakery kitchen this time.

- **The Oracle** (the model): The chef behind the pass. It reads whatever is in front of it, every turn, and cooks nothing itself.
- **The tickets** (the catalog): One per skill, on the rail over the pass: a name and a line on when to use it. They’re part of the system prompt, so they go with every request.
- **The shelf** (the skill bodies): One recipe book per skill, closed. The thicker the book, the more tokens it weighs. None of it reaches the model until someone fetches it.
- **The open book** (a loaded skill): `load_skill` returns the body as a tool result (the green bookmark). It joins the bandoneón like any result, so it lies open on the pass for the rest of the run.
- **The oven** (check_calendar): An ordinary tool. The recipe is what tells the model to use it.
- **The bar** (request size): How many tokens the next request carries. Its full length is what it would carry with all three books pasted into the system prompt.

In the EventBus panel, `load_skill` shows up as an ordinary `tool_execution_start` and `tool_execution_end`. The loop has no idea skills exist: to it, loading a manual is just another tool call.

## The code

**With astorlm:** Point `skillSources` at a folder of skills. The default `skillMode: 'on-demand'` puts the catalog in the system prompt and adds the `load_skill` tool for you.

**From scratch:** Read each `SKILL.md`, put one line per skill in the system prompt, and add a `load_skill` tool that returns a body. The loop from level 2 doesn’t change at all.

**With astorlm**

```ts
import { OpenAIProvider, createFileSystemSkillSource, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const checkCalendar = tool({
  name: 'check_calendar',
  description: 'Free baking slots and pickup times for a day, per production line.',
  schema: z.object({ day: z.string(), line: z.enum(['regular', 'nut-free']) }),
  execute: async ({ day, line }) => bakerySlots(day, line), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [checkCalendar],
  // Every folder in ./skills with a SKILL.md is one skill:
  //   skills/custom-cake-order/SKILL.md, skills/catering-quote/SKILL.md, …
  skillSources: [createFileSystemSkillSource({ dir: './skills' })],
  // The default. Only each skill's name and description go in the system prompt,
  // and the agent gets a load_skill tool to fetch a full body when it needs one.
  skillMode: 'on-demand',
  maxTurns: 10,
})

agent.on('tool-end', ({ name, output }) => {
  if (name === 'load_skill') console.log('loaded:', output.slice(0, 60))
})

const last = await agent.run('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.')
console.log(last.content)
```

**TypeScript**

```ts
// On-demand skills, from scratch. Plain fetch and node:fs, no SDK.
import { readdirSync, readFileSync } from 'node:fs'
import { join } from 'node:path'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// 1. Read the skills: one folder per skill, each with a SKILL.md.
//    The frontmatter (the block between the --- lines) holds name and description.
type Skill = { name: string; description: string; body: string }

function readSkill(file: string): Skill {
  const [, front = '', body = ''] = readFileSync(file, 'utf8').match(/^---\n([\s\S]*?)\n---\n?([\s\S]*)$/) ?? []
  const field = (key: string) => front.match(new RegExp(`^${key}:\\s*(.+)$`, 'm'))?.[1]?.trim() ?? ''
  return { name: field('name'), description: field('description'), body: body.trim() }
}

const SKILLS_DIR = './skills'
const skills = new Map(
  readdirSync(SKILLS_DIR, { withFileTypes: true })
    .filter((entry) => entry.isDirectory())
    .map((entry) => readSkill(join(SKILLS_DIR, entry.name, 'SKILL.md')))
    .map((skill) => [skill.name, skill]),
)

// 2. The catalog goes in the system prompt: one line per skill, never the body.
const catalog = [...skills.values()].map((s) => `- \`${s.name}\` — ${s.description}`).join('\n')
const SYSTEM = `You are the order assistant of a bakery.

<available-skills>
When the customer's request matches a skill, call load_skill with its name
before you answer, and follow the instructions it returns.
${catalog}
</available-skills>`

// 3. load_skill is a tool like any other. Its result is the body.
const loadSkill = ({ name }: { name: string }): string => {
  const skill = skills.get(name)
  if (!skill) throw new Error(`Unknown skill "${name}". Available: ${[...skills.keys()].join(', ')}`)
  return `<skill name="${skill.name}">\n${skill.body}\n</skill>`
}

type ToolFn = (args: Record<string, string>) => string | Promise<string>
const tools: Record<string, ToolFn> = {
  load_skill: (args) => loadSkill({ name: args.name ?? '' }),
  check_calendar: checkCalendar, // your code
}
const toolSchemas = [
  {
    type: 'function',
    function: {
      name: 'load_skill',
      description: "Load a skill's full instructions by name. Skills are listed in <available-skills>.",
      parameters: { type: 'object', properties: { name: { type: 'string' } }, required: ['name'] },
    },
  },
  /* check_calendar's schema */
]

// 4. The loop from level 2. Nothing in it knows about skills.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  const messages: Message[] = [
    { role: 'system', content: SYSTEM },
    { role: 'user', content: prompt },
  ]

  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      try {
        if (run) output = await run(JSON.parse(call.function.arguments))
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      // A loaded body lands here, in the history, and every later turn resends it.
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

console.log(await runAgent('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.'))
```

**Python**

```python
# On-demand skills, from scratch. Standard library only, no SDK.
import json
import re
import urllib.request
from pathlib import Path

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

# 1. Read the skills: one folder per skill, each with a SKILL.md.
#    The frontmatter (the block between the --- lines) holds name and description.
def read_skill(path):
    match = re.match(r"^---\n(.*?)\n---\n?(.*)$", path.read_text(encoding="utf-8"), re.S)
    front, body = match.groups() if match else ("", "")

    def field(key):
        found = re.search(rf"^{key}:\s*(.+)$", front, re.M)
        return found.group(1).strip() if found else ""

    return {"name": field("name"), "description": field("description"), "body": body.strip()}

SKILLS = {
    skill["name"]: skill
    for skill in (read_skill(folder / "SKILL.md") for folder in Path("skills").iterdir() if folder.is_dir())
}

# 2. The catalog goes in the system prompt: one line per skill, never the body.
CATALOG = "\n".join(f"- `{s['name']}` — {s['description']}" for s in SKILLS.values())
SYSTEM = f"""You are the order assistant of a bakery.

<available-skills>
When the customer's request matches a skill, call load_skill with its name
before you answer, and follow the instructions it returns.
{CATALOG}
</available-skills>"""

# 3. load_skill is a tool like any other. Its result is the body.
def load_skill(name):
    if name not in SKILLS:
        raise ValueError(f'Unknown skill "{name}". Available: {", ".join(SKILLS)}')
    return f'<skill name="{name}">\n{SKILLS[name]["body"]}\n</skill>'

TOOLS = {"load_skill": load_skill, "check_calendar": check_calendar}  # check_calendar: your code
TOOL_SCHEMAS = [
    {
        "type": "function",
        "function": {
            "name": "load_skill",
            "description": "Load a skill's full instructions by name. Skills are listed in <available-skills>.",
            "parameters": {"type": "object", "properties": {"name": {"type": "string"}}, "required": ["name"]},
        },
    },
    # check_calendar's schema
]

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

# 4. The loop from level 2. Nothing in it knows about skills.
def run_agent(prompt, max_turns=10):
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]

    for _ in range(max_turns):
        choice = chat(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            try:
                output = TOOLS[name](**json.loads(call["function"]["arguments"]))
            except Exception as err:
                output = f"Error: {err}"
            # A loaded body lands here, in the history, and every later turn resends it.
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

print(run_agent("I need a cake for 20 people this Saturday. Chocolate, and one guest can't have nuts."))
```

## What to watch

- **The description is the trigger.** The model picks a skill by its description alone, the same way it picks a tool (level 3). Say when to use it, not what it contains: “Use when a customer wants a cake baked for them” beats “Cake information”. In the animation, `allergen-info` also mentions allergies; only the descriptions keep the model from loading the wrong book.
- **Small models skip the step.** Some answer straight away without loading the skill they needed. Say it plainly in the system prompt (“load the matching skill before you answer”), try the filesystem mode, or, for a skill every request needs, keep it in `all`.
- **A loaded skill stays loaded.** Its body is in the history, so every later turn pays for it, and compaction (level 7) may truncate it later. Keep bodies short and split big skills in two.
- **Skills are instructions, so they are code.** Whoever writes a `SKILL.md` steers your agent. Don’t load skills from folders or registries you don’t control (level 14).

## Related patterns

- [0 · Your toolkit](https://harnesspatterns.dev/patterns/your-toolkit.md)
- [3 · Designing a tool](https://harnesspatterns.dev/patterns/designing-a-tool.md)
- [7 · The backpack fills up](https://harnesspatterns.dev/patterns/compaction.md)
- [9 · Memory](https://harnesspatterns.dev/patterns/memory.md)
- [14 · Security and sandboxing](https://harnesspatterns.dev/patterns/security.md)
