第 15 关
子智能体
- user
- assistant
- tool_result
EventBus
问题
很多请求里都藏着差事:筛二十条信息、读一个长页面、看四十条评价、在代码库里翻找一个函数。智能体需要的是每件差事的答案,而不是翻找的过程。
这就是 Hoarder,每件差事都亲自跑的智能体。每条搜索结果、每个页面都进了它的历史记录,循环在之后的每一轮都要把这些全部重新发送。等它回到你的问题时,它是在一堆只用了一分钟的信息底下读这个问题。请求变得很重,模型开始分心,再多一件差事就会把它推出窗口之外。
压缩(第 7 关)可以事后把这堆东西修剪掉。更好的办法是一开始就别堆起来。
解决方案
把差事交给子智能体。子智能体是一个完整的智能体,有自己的循环、自己的模型调用和自己的工具,而父智能体只把它看作一个工具。父智能体调用它时,一个全新的智能体会以空的历史记录和一条消息启动:那条消息就是父智能体写的简报。它去翻找资料、作答,然后被丢弃。父智能体只拿到那个答案,作为一个普通的工具结果。
-
亲力亲为
一个智能体,握着所有工具。每一次搜索、每一个页面,都由它自己来跑、自己来读。
每个原始结果都留在它的历史记录里,之后的每一轮都要重新发送。差事把问题埋了起来。
-
子智能体
把差事交给另一个以工具形式暴露的智能体。它从空白开始,拿到一份简报和自己的工具,然后用几行字作答。
父智能体保持小巧、专注。代价是:模型调用的总次数更多,而子智能体只知道简报里写的东西。
-
工作流
你的代码按照你写好的顺序调用各个智能体(第 1 关)。没有谁来决定委派:是你事先决定好的。
可预测,也容易推理。只有在请求到来之前你就知道所有步骤时,它才行得通。
还有两样东西是白送的。如果模型在同一条消息里申请两个子智能体,循环会像对待任何两个工具调用一样,把它们并行运行。而且每个子智能体可以有不同的 system prompt、更窄的工具集,甚至更小的模型:负责看评价的侦察员,没有任何理由去订东西。
编程智能体一直在用这一招:“探索一下这个仓库,告诉我认证在哪里处理”,就交给一个子智能体,它在五十个文件里 grep 一遍,带着三行字回来。
角色
还是那群熟悉的角色,这次换到了一家侦探事务所。
- 总部 父智能体
- 第 2 关的循环:Astor、神谕者和手风琴。它唯一的工具就是那两名侦察员。
- 电报机 子智能体工具
milonga_scout和food_scout在这里运行。简报作为工具的输入沿电线发下去;电报作为它的结果沿电线传上来。- 一扇外勤窗口 一次子智能体运行
- 一个完整的智能体:一名侦察员、他自己的神谕者、他自己的工具(那两家店铺)和他自己的手风琴。窗口的计数器就是它的上下文。它一作答,就消失了。
- 风箱褶的厚度 token
- 在这一关里,一褶有多厚,取决于它的消息有多重。一封三行的电报是薄薄一片。一页评价则是厚厚一块。
看看顶部的两根条。Parent 是父智能体请求真正的分量。All in 1 是如果父智能体亲自跑完两件差事、把每个页面都留在自己的历史记录里时的分量:它最后越过了压缩线。
事件日志里的 subagent 行是侦察员自己的事件。父智能体的 EventBus 从来看不到它们:它收到的只有每个子智能体工具的开始和结束。
代码
使用 astorlm:createSubagentTool 把一个提供商、一个 system prompt 和一组工具包装成父智能体可以调用的单个工具。每次调用都会启动一个新的子智能体,把简报跑到底,并返回它的最终文本。取消父智能体,子智能体也会一并取消。
从零手写:第 2 关的循环,改成把工具作为参数传入。子智能体就是一个工具,它的函数体会再次调用这个循环,带着新的消息和更少的工具。
import { OpenAIProvider, createLocalAgent, createSubagentTool, tool } from 'astorlm'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
const provider = new OpenAIProvider({ ...LLM, model: 'your-model' }) // e.g. 'llama3.1', 'gpt-4o-mini'
// The heavy tools: each one returns whole listings, pages or reviews.
const searchEvents = tool({
name: 'search_events',
description: 'Search tango events by neighborhood and date. Returns every match with its blurb.',
schema: z.object({ neighborhood: z.string(), date: z.string() }),
execute: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date), // your code
})
const readPage = tool({
name: 'read_page',
description: 'Read a web page and return its text.',
schema: z.object({ url: z.string() }),
execute: async ({ url }) => fetchText(url), // your code
})
const searchPlaces = tool({
name: 'search_places',
description: 'Search restaurants near a street, with their opening hours.',
schema: z.object({ near: z.string() }),
execute: async ({ near }) => placesApi.search(near), // your code
})
const readReviews = tool({
name: 'read_reviews',
description: 'Read the latest reviews of one restaurant.',
schema: z.object({ place: z.string() }),
execute: async ({ place }) => placesApi.reviews(place), // your code
})
// Each subagent is a whole agent, handed to the parent as ONE tool.
// It gets its own system prompt, only the tools it needs, and a fresh history on every call.
const milongaScout = createSubagentTool({
name: 'milonga_scout',
description: 'Finds tango events. Give it a full brief: it knows nothing else about the conversation.',
provider, // could be a smaller, cheaper model
systemPrompt: 'You find milongas in Buenos Aires. Reply in 3 lines: name, address, times. No lists, no links.',
tools: [searchEvents, readPage],
maxTurns: 6,
})
const foodScout = createSubagentTool({
name: 'food_scout',
description: 'Finds places to eat. Give it a full brief: it knows nothing else about the conversation.',
provider,
systemPrompt: 'You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.',
tools: [searchPlaces, readReviews],
maxTurns: 6,
})
// The parent only sees two tools. It never gets the listings, pages or reviews: just each scout's final text.
const agent = await createLocalAgent({
provider,
systemPrompt: 'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
tools: [milongaScout, foodScout],
maxTurns: 6,
})
const answer = await agent.run('I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.')
console.log(answer.content)
// Both scouts were asked for in one message, so the loop ran them in parallel.
// Cancelling the parent (abortSignal) cancels any scout still out.
// Subagents, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type Args = Record<string, string>
type ToolFn = (args: Args) => Promise<string>
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 1. The loop from level 2, with its tools passed in. `messages` is born and dies inside each call.
async function runAgent(system: string, prompt: string, tools: Record<string, ToolFn>, schemas: object[], maxTurns = 6): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: system },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: schemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
// Run every call of this message at once, and add the results in order.
const calls = reply.tool_calls ?? []
const outputs = await Promise.all(
calls.map(async (call) => {
try {
const run = tools[call.function.name]
return run ? await run(JSON.parse(call.function.arguments)) : `Unknown tool: ${call.function.name}`
} catch (err) {
return `Error: ${err instanceof Error ? err.message : err}`
}
}),
)
calls.forEach((call, i) => messages.push({ role: 'tool', tool_call_id: call.id, content: outputs[i]! }))
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// 2. The scouts' own tools: the heavy ones. Your code.
const scoutTools: Record<string, ToolFn> = {
search_events: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date),
read_page: async ({ url }) => fetchText(url),
search_places: async ({ near }) => placesApi.search(near),
read_reviews: async ({ place }) => placesApi.reviews(place),
}
const pick = (...names: string[]) => Object.fromEntries(names.map((name) => [name, scoutTools[name]!]))
// 3. A subagent is a tool whose body is another runAgent call: new messages, fewer tools, its own prompt.
// Only its final text comes back. Everything it read dies with its `messages`.
const parentTools: Record<string, ToolFn> = {
milonga_scout: ({ task }) =>
runAgent('You find milongas in Buenos Aires. Reply in 3 lines: name, address, times.', task, pick('search_events', 'read_page'), [/* their schemas */]),
food_scout: ({ task }) =>
runAgent('You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.', task, pick('search_places', 'read_reviews'), [/* their schemas */]),
}
// Both take one string, `task`. The description tells the parent to write a full brief.
const parentSchemas = [/* milonga_scout(task), food_scout(task) */]
// 4. The parent: the same loop, and all it ever sees of the scouts is two short answers.
const answer = await runAgent(
'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
'I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.',
parentTools,
parentSchemas,
)
console.log(answer)
# Subagents, from scratch. Standard library only, no SDK.
import json
import urllib.request
from concurrent.futures import ThreadPoolExecutor
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
def call_tool(tools, call):
try:
return tools[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
return f"Error: {err}"
# 1. The loop from level 2, with its tools passed in. `messages` is born and dies inside each call.
def run_agent(system, prompt, tools, schemas, max_turns=6):
messages = [{"role": "system", "content": system}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": schemas})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
# Run every call of this message at once, and add the results in order.
calls = reply.get("tool_calls", [])
with ThreadPoolExecutor() as pool:
outputs = list(pool.map(lambda call: call_tool(tools, call), calls))
for call, output in zip(calls, outputs):
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# 2. The scouts' own tools: the heavy ones. Your code.
SCOUT_TOOLS = {
"search_events": lambda neighborhood, date: events_api.search(neighborhood, date),
"read_page": lambda url: fetch_text(url),
"search_places": lambda near: places_api.search(near),
"read_reviews": lambda place: places_api.reviews(place),
}
def pick(*names):
return {name: SCOUT_TOOLS[name] for name in names}
# 3. A subagent is a tool whose body is another run_agent call: new messages, fewer tools, its own prompt.
# Only its final text comes back. Everything it read dies with its `messages`.
def milonga_scout(task):
system = "You find milongas in Buenos Aires. Reply in 3 lines: name, address, times."
return run_agent(system, task, pick("search_events", "read_page"), [...]) # their schemas
def food_scout(task):
system = "You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why."
return run_agent(system, task, pick("search_places", "read_reviews"), [...]) # their schemas
# Both take one string, `task`. The description tells the parent to write a full brief.
PARENT_TOOLS = {"milonga_scout": milonga_scout, "food_scout": food_scout}
PARENT_SCHEMAS = [...] # milonga_scout(task), food_scout(task)
# 4. The parent: the same loop, and all it ever sees of the scouts is two short answers.
answer = run_agent(
"You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.",
"I'm staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.",
PARENT_TOOLS,
PARENT_SCHEMAS,
)
print(answer)
注意事项
- 简报就是它知道的全部。子智能体从没见过这段对话。“找我们刚才说的那个”对它毫无意义。在工具描述里要求父智能体写一份完整的简报:目标、约束条件,以及好答案应该是什么样子。
- 要求简短、固定的格式。整件事的意义就在于得到一个小结果。像“用 3 行回答:名称、地址、时间”这样的 system prompt,能防止侦察员把它那堆东西又贴回给父智能体。
- 省下的是上下文,不是钱。翻找照样在发生,只是发生在另一个智能体的请求里。总成本往往更高。当父智能体的专注值得这个价钱时才用子智能体,差事允许的话就给它们换个更便宜的模型。
- 只拆分彼此独立的部分。两名侦察员能并排跑,是因为谁都不需要对方。如果第二件差事需要第一件的答案,就一个接一个地调用,或者干脆留在一个智能体里做。
- 限定它的工具,并限制嵌套深度。只给每个子智能体它那件差事需要的工具;给它配属于它自己的子智能体之前,要三思。每多一层,调用次数就成倍增加,而深处的一个失败,传上来时只剩一行让人摸不着头脑的话。