Skip to content
astorlm
← Map

Level 15

Subagents

Some errands are heavy and self-contained. Hand them to another agent: it starts with a clean history, does the digging, and sends back only what you need.
1/22 Bandoneón folds:
  • user
  • assistant
  • tool_result
A detective agency. Upstairs, the Oracle and Astor at headquarters. Downstairs, two empty field windows: nobody is out on a case yet.

EventBus

The problem

Plenty of requests hide errands inside them: search twenty listings, read a long page, go through forty reviews, dig through a codebase to find one function. The agent needs the answer to each errand. It doesn’t need the digging.

That’s Hoarder, the agent that runs every errand itself. Each search result and each page joins its history, and the loop resends all of it on every turn after. By the time it gets back to your question, it’s reading it under a pile of listings it only needed for a minute. The requests are heavy, the model gets distracted, and one more errand pushes it past the window.

Compaction (level 7) can trim the pile afterwards. Better not to build it in the first place.

The solution

Give the errand to a subagent. A subagent is a whole agent, with its own loop, its own model calls and its own tools, that the parent sees as one tool. When the parent calls it, a fresh agent starts with an empty history and one message: the brief the parent wrote. It does the digging, answers, and is thrown away. The parent only gets that answer, as an ordinary tool result.

  • Do it yourself

    One agent, every tool. It runs each search and reads each page itself.

    Every raw result stays in its history, and every later turn resends it. The errands bury the question.

  • Subagent

    Hand the errand to another agent, exposed as a tool. It starts empty, gets a brief and its own tools, and answers in a few lines.

    The parent stays small and focused. The price: more model calls in total, and the subagent only knows what the brief says.

  • Workflow

    Your code calls the agents, in an order you wrote (level 1). Nobody decides to delegate: you did, ahead of time.

    Predictable and cheap to reason about. It only works when you know the steps before the request arrives.

Two things come for free. If the model asks for two subagents in the same message, the loop runs them in parallel, like any two tool calls. And each subagent can get a different system prompt, a narrower set of tools, even a smaller model: the scout that reads reviews has no reason to book anything.

Coding agents lean on this all the time: “explore the repo and tell me where auth is handled” goes to a subagent that greps through fifty files and comes back with three lines.

The cast

Same cast as always, at a detective agency this time.

Headquarters the parent agent
The loop from level 2: Astor, the Oracle and the bandoneón. Its only tools are the two scouts.
The telegraph subagent tools
Where milonga_scout and food_scout run. A brief goes down the wire as the tool’s input; a telegram comes back up as its result.
A field window one subagent run
A whole agent: a scout, his own Oracle, his own tools (the two shops) and his own bandoneón. The window’s counter is its context. When it answers, it’s gone.
Fold thickness tokens
In this level a fold is as thick as its message is heavy. A three-line telegram is a sliver. A page of reviews is a slab.

Watch the two bars at the top. Parent is what the parent’s request really weighs. All in 1 is what it would weigh if the parent had run both errands itself, with every page in its own history: it ends past the compaction line.

The subagent lines in the event log are the scouts’ own events. The parent’s EventBus never sees them: all it gets is each subagent tool’s start and end.

The code

With astorlm: createSubagentTool wraps a provider, a system prompt and a set of tools into one tool the parent can call. Every call starts a new child agent, runs the brief to the end and returns its final text. Cancelling the parent cancels the child.

From scratch: The loop from level 2, taking its tools as an argument. A subagent is a tool whose body calls that loop again, with new messages and fewer tools.

import { OpenAIProvider, createLocalAgent, createSubagentTool, tool } from 'astorlm'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
const provider = new OpenAIProvider({ ...LLM, model: 'your-model' }) // e.g. 'llama3.1', 'gpt-4o-mini'

// The heavy tools: each one returns whole listings, pages or reviews.
const searchEvents = tool({
  name: 'search_events',
  description: 'Search tango events by neighborhood and date. Returns every match with its blurb.',
  schema: z.object({ neighborhood: z.string(), date: z.string() }),
  execute: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date), // your code
})
const readPage = tool({
  name: 'read_page',
  description: 'Read a web page and return its text.',
  schema: z.object({ url: z.string() }),
  execute: async ({ url }) => fetchText(url), // your code
})
const searchPlaces = tool({
  name: 'search_places',
  description: 'Search restaurants near a street, with their opening hours.',
  schema: z.object({ near: z.string() }),
  execute: async ({ near }) => placesApi.search(near), // your code
})
const readReviews = tool({
  name: 'read_reviews',
  description: 'Read the latest reviews of one restaurant.',
  schema: z.object({ place: z.string() }),
  execute: async ({ place }) => placesApi.reviews(place), // your code
})

// Each subagent is a whole agent, handed to the parent as ONE tool.
// It gets its own system prompt, only the tools it needs, and a fresh history on every call.
const milongaScout = createSubagentTool({
  name: 'milonga_scout',
  description: 'Finds tango events. Give it a full brief: it knows nothing else about the conversation.',
  provider, // could be a smaller, cheaper model
  systemPrompt: 'You find milongas in Buenos Aires. Reply in 3 lines: name, address, times. No lists, no links.',
  tools: [searchEvents, readPage],
  maxTurns: 6,
})
const foodScout = createSubagentTool({
  name: 'food_scout',
  description: 'Finds places to eat. Give it a full brief: it knows nothing else about the conversation.',
  provider,
  systemPrompt: 'You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.',
  tools: [searchPlaces, readReviews],
  maxTurns: 6,
})

// The parent only sees two tools. It never gets the listings, pages or reviews: just each scout's final text.
const agent = await createLocalAgent({
  provider,
  systemPrompt: 'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
  tools: [milongaScout, foodScout],
  maxTurns: 6,
})

const answer = await agent.run('I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.')
console.log(answer.content)
// Both scouts were asked for in one message, so the loop ran them in parallel.
// Cancelling the parent (abortSignal) cancels any scout still out.

What to watch

  • The brief is all it knows. The subagent never saw the conversation. “Find the one we talked about” means nothing to it. Tell the parent, in the tool’s description, to write a full brief: the goal, the constraints, and what a good answer looks like.
  • Ask for a short, fixed shape. The whole point is a small result. A system prompt like “reply in 3 lines: name, address, times” keeps the scout from pasting its pile back into the parent.
  • It saves context, not money. The digging still happens, in another agent’s requests. Often it costs more in total. Use subagents when the parent’s focus is worth it, and hand them a cheaper model when the errand allows.
  • Only split what’s independent. Two scouts can run side by side because neither needs the other. If the second errand needs the first one’s answer, call them one after the other, or keep it in one agent.
  • Scope its tools, and cap the depth. Give each subagent only the tools its errand needs, and think twice before giving it subagents of its own. Each level multiplies the calls, and a failure deep down arrives as one confusing line.