Skip to content
astorlm
← Map

Level 8

On-demand skills

A skill is a manual for one kind of job. The agent always sees the list of manuals, and reads one only when a request calls for it.
1/16 Bandoneón folds:
  • user
  • assistant
  • tool_result
A bakery’s order assistant. Three recipe books on the shelf: those are its skills. Over the pass hangs one ticket per book, with its name and a line on when to use it. Only the tickets go in the system prompt.

EventBus

The problem

A bakery’s order assistant needs to know the house rules. Which cake size feeds 20 people. That nut-free cakes only bake on Friday mornings. How much the deposit is. How to quote a catering tray. What goes into each product already on the shelves. The model knows none of it.

The obvious fix is to write it all down and paste it into the system prompt. It works on day one. Then the manuals grow, and there are ten of them. Every request now carries every manual, on every turn, whether the customer wants a wedding cake or just the opening hours. You pay for all of it, the requests get slower, and the one rule that matters is buried among the ones that don’t.

That’s Tomebloat, the bloated system prompt. It feeds Gulp from level 7: a prompt that’s already full leaves less room for the conversation itself.

The solution

Split each manual in two. A description of one or two lines says when the skill applies. The body says how to do the job. Only the descriptions go in the system prompt, as a catalog. When a request matches one, the model asks for its body, and the body joins the conversation from then on.

That’s progressive disclosure: show the index first, and the detail only when it’s needed. The usual format is a folder per skill with a SKILL.md file. The frontmatter at the top (the block between the --- lines) holds the name and the description. Everything below it is the body.

---
name: custom-cake-order
description: Cakes made to order. Use when a customer wants a cake baked for them — size by guests, flavors, allergies, lead time, deposit.
---

# Custom cake orders

1. Size by guests: up to 12 → 18 cm, up to 24 → 24 cm, more → two tiers.
2. Flavors: chocolate, vanilla, dulce de leche. Nothing else.
3. Nut allergy → the nut-free line. It only bakes on Friday mornings.
4. Check the calendar before you promise a date. Never less than 72 hours.
5. Quote the 50% deposit, and ask before you book. Never book on your own.

There are three common ways to hand skills to the model. astorlm supports all three with skillMode:

  • all

    Every skill’s full body goes into the system prompt, from the first request.

    No extra calls, and the model can’t skip a manual. Fine for two or three short skills that almost every request needs. Past that, it’s Tomebloat.

  • on-demand

    The system prompt lists each skill’s name and description. A load_skill tool returns a full body when the model asks for it.

    Scales to dozens of skills. Costs one tool call per skill loaded, and small models sometimes answer without loading the skill they needed.

  • filesystem

    The catalog also gives each SKILL.md’s path, and the model opens it with its ordinary file-reading tool.

    The same idea with no special tool. It’s how Claude Code, Codex and Gemini CLI do it, so the same skills folder works in all of them. Needs an agent that can read files.

A skill is not a tool. A tool is something the loop runs for the model. A skill is something the model reads, and it can tell the model which tools to call and in what order, like the recipe that says “check the calendar before you promise a date”.

The cast

Same cast as always, in a bakery kitchen this time.

The Oracle the model
The chef behind the pass. It reads whatever is in front of it, every turn, and cooks nothing itself.
The tickets the catalog
One per skill, on the rail over the pass: a name and a line on when to use it. They’re part of the system prompt, so they go with every request.
The shelf the skill bodies
One recipe book per skill, closed. The thicker the book, the more tokens it weighs. None of it reaches the model until someone fetches it.
The open book a loaded skill
load_skill returns the body as a tool result (the green bookmark). It joins the bandoneón like any result, so it lies open on the pass for the rest of the run.
The oven check_calendar
An ordinary tool. The recipe is what tells the model to use it.
The bar request size
How many tokens the next request carries. Its full length is what it would carry with all three books pasted into the system prompt.

In the EventBus panel, load_skill shows up as an ordinary tool_execution_start and tool_execution_end. The loop has no idea skills exist: to it, loading a manual is just another tool call.

The code

With astorlm: Point skillSources at a folder of skills. The default skillMode: 'on-demand' puts the catalog in the system prompt and adds the load_skill tool for you.

From scratch: Read each SKILL.md, put one line per skill in the system prompt, and add a load_skill tool that returns a body. The loop from level 2 doesn’t change at all.

import { OpenAIProvider, createFileSystemSkillSource, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const checkCalendar = tool({
  name: 'check_calendar',
  description: 'Free baking slots and pickup times for a day, per production line.',
  schema: z.object({ day: z.string(), line: z.enum(['regular', 'nut-free']) }),
  execute: async ({ day, line }) => bakerySlots(day, line), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [checkCalendar],
  // Every folder in ./skills with a SKILL.md is one skill:
  //   skills/custom-cake-order/SKILL.md, skills/catering-quote/SKILL.md, …
  skillSources: [createFileSystemSkillSource({ dir: './skills' })],
  // The default. Only each skill's name and description go in the system prompt,
  // and the agent gets a load_skill tool to fetch a full body when it needs one.
  skillMode: 'on-demand',
  maxTurns: 10,
})

agent.on('tool-end', ({ name, output }) => {
  if (name === 'load_skill') console.log('loaded:', output.slice(0, 60))
})

const last = await agent.run('I need a cake for 20 people this Saturday. Chocolate, and one guest can’t have nuts.')
console.log(last.content)

What to watch

  • The description is the trigger. The model picks a skill by its description alone, the same way it picks a tool (level 3). Say when to use it, not what it contains: “Use when a customer wants a cake baked for them” beats “Cake information”. In the animation, allergen-info also mentions allergies; only the descriptions keep the model from loading the wrong book.
  • Small models skip the step. Some answer straight away without loading the skill they needed. Say it plainly in the system prompt (“load the matching skill before you answer”), try the filesystem mode, or, for a skill every request needs, keep it in all.
  • A loaded skill stays loaded. Its body is in the history, so every later turn pays for it, and compaction (level 7) may truncate it later. Keep bodies short and split big skills in two.
  • Skills are instructions, so they are code. Whoever writes a SKILL.md steers your agent. Don’t load skills from folders or registries you don’t control (level 14).