Nível 15
Subagentes
- user
- assistant
- tool_result
EventBus
O problema
Muitas requisições escondem tarefas dentro delas: vasculhar vinte anúncios, ler uma página longa, passar por quarenta avaliações, percorrer uma base de código para achar uma função. O agente precisa da resposta de cada tarefa. Ele não precisa do trabalho pesado de busca.
Esse é o Hoarder, o agente que faz cada tarefa sozinho. Cada resultado de busca e cada página entram no histórico dele, e o loop reenvia tudo isso a cada turno seguinte. Quando ele volta à sua pergunta, já a está lendo debaixo de uma pilha de anúncios de que só precisou por um minuto. As requisições ficam pesadas, o modelo se distrai, e mais uma tarefa o empurra para além da janela.
A compactação (nível 7) pode aparar a pilha depois. Melhor nem montá-la.
A solução
Dê a tarefa a um subagente. Um subagente é um agente completo, com o próprio loop, as próprias chamadas ao modelo e as próprias ferramentas, que o agente pai enxerga como uma ferramenta. Quando o pai a chama, um agente novo começa com o histórico vazio e uma mensagem: o briefing que o pai escreveu. Ele faz o trabalho pesado, responde e é descartado. O pai só recebe essa resposta, como um resultado de ferramenta comum.
-
Fazer sozinho
Um agente, todas as ferramentas. Ele roda cada busca e lê cada página por conta própria.
Cada resultado bruto fica no histórico dele, e cada turno seguinte o reenvia. As tarefas soterram a pergunta.
-
Subagente
Passar a tarefa para outro agente, exposto como ferramenta. Ele começa vazio, recebe um briefing e as próprias ferramentas, e responde em poucas linhas.
O agente pai fica pequeno e focado. O preço: mais chamadas ao modelo no total, e o subagente só sabe o que o briefing diz.
-
Workflow
O seu código chama os agentes, numa ordem que você escreveu (nível 1). Ninguém decide delegar: você decidiu, de antemão.
Previsível e fácil de entender. Só funciona quando você conhece os passos antes de a requisição chegar.
Duas coisas vêm de brinde. Se o modelo pede dois subagentes na mesma mensagem, o loop os roda em paralelo, como quaisquer duas chamadas de ferramenta. E cada subagente pode ter um system prompt diferente, um conjunto mais restrito de ferramentas, até um modelo menor: o investigador que lê avaliações não tem por que reservar nada.
Agentes de código se apoiam nisso o tempo todo: “explore o repositório e me diga onde a autenticação é tratada” vai para um subagente que faz grep em cinquenta arquivos e volta com três linhas.
O elenco
O mesmo elenco de sempre, desta vez numa agência de detetives.
- A central o agente pai
- O loop do nível 2: Astor, o Oráculo e o bandoneón. As únicas ferramentas dele são os dois investigadores.
- O telégrafo ferramentas de subagentes
-
Onde rodam
milonga_scoutefood_scout. Um briefing desce pelo fio como entrada da ferramenta; um telegrama sobe como resultado dela. - Uma janela de campo uma execução de subagente
- Um agente completo: um investigador, o próprio Oráculo, as próprias ferramentas (as duas lojas) e o próprio bandoneón. O contador da janela é o contexto dele. Quando ele responde, some.
- A espessura das dobras tokens
- Neste nível, uma dobra é tão grossa quanto a mensagem é pesada. Um telegrama de três linhas é uma lasca. Uma página de avaliações é um bloco.
Observe as duas barras no topo. Parent é o quanto a requisição do pai pesa de verdade. All in 1 é o quanto pesaria se o pai tivesse feito as duas tarefas sozinho, com cada página no próprio histórico: ela termina acima da linha de compactação.
As linhas subagent no log de eventos são os eventos dos próprios investigadores. O EventBus do pai
nunca as vê: tudo o que ele recebe é o início e o fim de cada ferramenta de subagente.
O código
Com astorlm: createSubagentTool embrulha um provedor, um system prompt e um conjunto
de ferramentas numa única ferramenta que o pai pode chamar. Cada chamada inicia um novo agente filho, roda o
briefing até o fim e devolve o texto final dele. Cancelar o pai cancela o filho.
Do zero: O loop do nível 2, recebendo as ferramentas como argumento. Um subagente é uma ferramenta cujo corpo chama esse loop de novo, com mensagens novas e menos ferramentas.
import { OpenAIProvider, createLocalAgent, createSubagentTool, tool } from 'astorlm'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
const provider = new OpenAIProvider({ ...LLM, model: 'your-model' }) // e.g. 'llama3.1', 'gpt-4o-mini'
// The heavy tools: each one returns whole listings, pages or reviews.
const searchEvents = tool({
name: 'search_events',
description: 'Search tango events by neighborhood and date. Returns every match with its blurb.',
schema: z.object({ neighborhood: z.string(), date: z.string() }),
execute: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date), // your code
})
const readPage = tool({
name: 'read_page',
description: 'Read a web page and return its text.',
schema: z.object({ url: z.string() }),
execute: async ({ url }) => fetchText(url), // your code
})
const searchPlaces = tool({
name: 'search_places',
description: 'Search restaurants near a street, with their opening hours.',
schema: z.object({ near: z.string() }),
execute: async ({ near }) => placesApi.search(near), // your code
})
const readReviews = tool({
name: 'read_reviews',
description: 'Read the latest reviews of one restaurant.',
schema: z.object({ place: z.string() }),
execute: async ({ place }) => placesApi.reviews(place), // your code
})
// Each subagent is a whole agent, handed to the parent as ONE tool.
// It gets its own system prompt, only the tools it needs, and a fresh history on every call.
const milongaScout = createSubagentTool({
name: 'milonga_scout',
description: 'Finds tango events. Give it a full brief: it knows nothing else about the conversation.',
provider, // could be a smaller, cheaper model
systemPrompt: 'You find milongas in Buenos Aires. Reply in 3 lines: name, address, times. No lists, no links.',
tools: [searchEvents, readPage],
maxTurns: 6,
})
const foodScout = createSubagentTool({
name: 'food_scout',
description: 'Finds places to eat. Give it a full brief: it knows nothing else about the conversation.',
provider,
systemPrompt: 'You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.',
tools: [searchPlaces, readReviews],
maxTurns: 6,
})
// The parent only sees two tools. It never gets the listings, pages or reviews: just each scout's final text.
const agent = await createLocalAgent({
provider,
systemPrompt: 'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
tools: [milongaScout, foodScout],
maxTurns: 6,
})
const answer = await agent.run('I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.')
console.log(answer.content)
// Both scouts were asked for in one message, so the loop ran them in parallel.
// Cancelling the parent (abortSignal) cancels any scout still out.
// Subagents, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type Args = Record<string, string>
type ToolFn = (args: Args) => Promise<string>
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 1. The loop from level 2, with its tools passed in. `messages` is born and dies inside each call.
async function runAgent(system: string, prompt: string, tools: Record<string, ToolFn>, schemas: object[], maxTurns = 6): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: system },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: schemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
// Run every call of this message at once, and add the results in order.
const calls = reply.tool_calls ?? []
const outputs = await Promise.all(
calls.map(async (call) => {
try {
const run = tools[call.function.name]
return run ? await run(JSON.parse(call.function.arguments)) : `Unknown tool: ${call.function.name}`
} catch (err) {
return `Error: ${err instanceof Error ? err.message : err}`
}
}),
)
calls.forEach((call, i) => messages.push({ role: 'tool', tool_call_id: call.id, content: outputs[i]! }))
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// 2. The scouts' own tools: the heavy ones. Your code.
const scoutTools: Record<string, ToolFn> = {
search_events: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date),
read_page: async ({ url }) => fetchText(url),
search_places: async ({ near }) => placesApi.search(near),
read_reviews: async ({ place }) => placesApi.reviews(place),
}
const pick = (...names: string[]) => Object.fromEntries(names.map((name) => [name, scoutTools[name]!]))
// 3. A subagent is a tool whose body is another runAgent call: new messages, fewer tools, its own prompt.
// Only its final text comes back. Everything it read dies with its `messages`.
const parentTools: Record<string, ToolFn> = {
milonga_scout: ({ task }) =>
runAgent('You find milongas in Buenos Aires. Reply in 3 lines: name, address, times.', task, pick('search_events', 'read_page'), [/* their schemas */]),
food_scout: ({ task }) =>
runAgent('You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.', task, pick('search_places', 'read_reviews'), [/* their schemas */]),
}
// Both take one string, `task`. The description tells the parent to write a full brief.
const parentSchemas = [/* milonga_scout(task), food_scout(task) */]
// 4. The parent: the same loop, and all it ever sees of the scouts is two short answers.
const answer = await runAgent(
'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
'I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.',
parentTools,
parentSchemas,
)
console.log(answer)
# Subagents, from scratch. Standard library only, no SDK.
import json
import urllib.request
from concurrent.futures import ThreadPoolExecutor
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
def call_tool(tools, call):
try:
return tools[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
return f"Error: {err}"
# 1. The loop from level 2, with its tools passed in. `messages` is born and dies inside each call.
def run_agent(system, prompt, tools, schemas, max_turns=6):
messages = [{"role": "system", "content": system}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": schemas})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
# Run every call of this message at once, and add the results in order.
calls = reply.get("tool_calls", [])
with ThreadPoolExecutor() as pool:
outputs = list(pool.map(lambda call: call_tool(tools, call), calls))
for call, output in zip(calls, outputs):
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# 2. The scouts' own tools: the heavy ones. Your code.
SCOUT_TOOLS = {
"search_events": lambda neighborhood, date: events_api.search(neighborhood, date),
"read_page": lambda url: fetch_text(url),
"search_places": lambda near: places_api.search(near),
"read_reviews": lambda place: places_api.reviews(place),
}
def pick(*names):
return {name: SCOUT_TOOLS[name] for name in names}
# 3. A subagent is a tool whose body is another run_agent call: new messages, fewer tools, its own prompt.
# Only its final text comes back. Everything it read dies with its `messages`.
def milonga_scout(task):
system = "You find milongas in Buenos Aires. Reply in 3 lines: name, address, times."
return run_agent(system, task, pick("search_events", "read_page"), [...]) # their schemas
def food_scout(task):
system = "You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why."
return run_agent(system, task, pick("search_places", "read_reviews"), [...]) # their schemas
# Both take one string, `task`. The description tells the parent to write a full brief.
PARENT_TOOLS = {"milonga_scout": milonga_scout, "food_scout": food_scout}
PARENT_SCHEMAS = [...] # milonga_scout(task), food_scout(task)
# 4. The parent: the same loop, and all it ever sees of the scouts is two short answers.
answer = run_agent(
"You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.",
"I'm staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.",
PARENT_TOOLS,
PARENT_SCHEMAS,
)
print(answer)
O que observar
- O briefing é tudo o que ele sabe. O subagente nunca viu a conversa. “Procure aquele de que falamos” não significa nada para ele. Diga ao pai, na descrição da ferramenta, para escrever um briefing completo: o objetivo, as restrições e como é uma boa resposta.
- Peça um formato curto e fixo. A ideia toda é um resultado pequeno. Um system prompt como “responda em 3 linhas: nome, endereço, horários” impede o investigador de colar a pilha dele de volta no pai.
- Economiza contexto, não dinheiro. O trabalho pesado continua acontecendo, nas requisições de outro agente. Muitas vezes custa mais no total. Use subagentes quando o foco do pai valer a pena, e dê a eles um modelo mais barato quando a tarefa permitir.
- Divida só o que é independente. Dois investigadores podem rodar lado a lado porque nenhum precisa do outro. Se a segunda tarefa precisa da resposta da primeira, chame um depois do outro, ou mantenha tudo num só agente.
- Restrinja as ferramentas, e limite a profundidade. Dê a cada subagente só as ferramentas de que a tarefa precisa, e pense duas vezes antes de dar a ele subagentes próprios. Cada nível multiplica as chamadas, e uma falha lá no fundo chega como uma única linha confusa.