> Agent Harness Patterns 第 15 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/subagents · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 子智能体

有些差事既繁重又自成一体。把它们交给另一个智能体：它从干净的历史记录开始，自己去翻找资料，只把你需要的东西送回来。

## 问题

很多请求里都藏着差事：筛二十条信息、读一个长页面、看四十条评价、在代码库里翻找一个函数。智能体需要的是每件差事的*答案*，而不是翻找的过程。

这就是 Hoarder，每件差事都亲自跑的智能体。每条搜索结果、每个页面都进了它的历史记录，循环在之后的每一轮都要把这些全部重新发送。等它回到你的问题时，它是在一堆只用了一分钟的信息底下读这个问题。请求变得很重，模型开始分心，再多一件差事就会把它推出窗口之外。

压缩（第 7 关）可以事后把这堆东西修剪掉。更好的办法是一开始就别堆起来。

## 解决方案

**把差事交给子智能体。**子智能体是一个完整的智能体，有自己的循环、自己的模型调用和自己的工具，而父智能体只把它看作*一个工具*。父智能体调用它时，一个全新的智能体会以空的历史记录和一条消息启动：那条消息就是父智能体写的简报。它去翻找资料、作答，然后被丢弃。父智能体只拿到那个答案，作为一个普通的工具结果。

- **亲力亲为**
   一个智能体，握着所有工具。每一次搜索、每一个页面，都由它自己来跑、自己来读。
   每个原始结果都留在它的历史记录里，之后的每一轮都要重新发送。差事把问题埋了起来。
- **子智能体**
   把差事交给另一个以工具形式暴露的智能体。它从空白开始，拿到一份简报和自己的工具，然后用几行字作答。
   父智能体保持小巧、专注。代价是：模型调用的总次数更多，而子智能体只知道简报里写的东西。
- **工作流**
   你的代码按照你写好的顺序调用各个智能体（第 1 关）。没有谁来决定委派：是你事先决定好的。
   可预测，也容易推理。只有在请求到来之前你就知道所有步骤时，它才行得通。

还有两样东西是白送的。如果模型在同一条消息里申请两个子智能体，循环会像对待任何两个工具调用一样，把它们**并行**运行。而且每个子智能体可以有不同的 system prompt、更窄的工具集，甚至更小的模型：负责看评价的侦察员，没有任何理由去订东西。

编程智能体一直在用这一招：“探索一下这个仓库，告诉我认证在哪里处理”，就交给一个子智能体，它在五十个文件里 grep 一遍，带着三行字回来。

## 角色

还是那群熟悉的角色，这次换到了一家侦探事务所。

- **总部** (父智能体): 第 2 关的循环：Astor、神谕者和手风琴。它唯一的工具就是那两名侦察员。
- **电报机** (子智能体工具): `milonga_scout` 和 `food_scout` 在这里运行。简报作为工具的输入沿电线发下去；电报作为它的结果沿电线传上来。
- **一扇外勤窗口** (一次子智能体运行): 一个完整的智能体：一名侦察员、他自己的神谕者、他自己的工具（那两家店铺）和他自己的手风琴。窗口的计数器就是它的上下文。它一作答，就消失了。
- **风箱褶的厚度** (token): 在这一关里，一褶有多厚，取决于它的消息有多重。一封三行的电报是薄薄一片。一页评价则是厚厚一块。

看看顶部的两根条。*Parent* 是父智能体请求真正的分量。*All in 1* 是如果父智能体亲自跑完两件差事、把每个页面都留在自己的历史记录里时的分量：它最后越过了压缩线。

事件日志里的 `subagent` 行是侦察员自己的事件。父智能体的 EventBus 从来看不到它们：它收到的只有每个子智能体工具的开始和结束。

## 代码

**使用 astorlm：**`createSubagentTool` 把一个提供商、一个 system prompt 和一组工具包装成父智能体可以调用的单个工具。每次调用都会启动一个新的子智能体，把简报跑到底，并返回它的最终文本。取消父智能体，子智能体也会一并取消。

**从零手写：**第 2 关的循环，改成把工具作为参数传入。子智能体就是一个工具，它的函数体会再次调用这个循环，带着新的消息和更少的工具。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, createSubagentTool, tool } from 'astorlm'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
const provider = new OpenAIProvider({ ...LLM, model: 'your-model' }) // e.g. 'llama3.1', 'gpt-4o-mini'

// The heavy tools: each one returns whole listings, pages or reviews.
const searchEvents = tool({
  name: 'search_events',
  description: 'Search tango events by neighborhood and date. Returns every match with its blurb.',
  schema: z.object({ neighborhood: z.string(), date: z.string() }),
  execute: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date), // your code
})
const readPage = tool({
  name: 'read_page',
  description: 'Read a web page and return its text.',
  schema: z.object({ url: z.string() }),
  execute: async ({ url }) => fetchText(url), // your code
})
const searchPlaces = tool({
  name: 'search_places',
  description: 'Search restaurants near a street, with their opening hours.',
  schema: z.object({ near: z.string() }),
  execute: async ({ near }) => placesApi.search(near), // your code
})
const readReviews = tool({
  name: 'read_reviews',
  description: 'Read the latest reviews of one restaurant.',
  schema: z.object({ place: z.string() }),
  execute: async ({ place }) => placesApi.reviews(place), // your code
})

// Each subagent is a whole agent, handed to the parent as ONE tool.
// It gets its own system prompt, only the tools it needs, and a fresh history on every call.
const milongaScout = createSubagentTool({
  name: 'milonga_scout',
  description: 'Finds tango events. Give it a full brief: it knows nothing else about the conversation.',
  provider, // could be a smaller, cheaper model
  systemPrompt: 'You find milongas in Buenos Aires. Reply in 3 lines: name, address, times. No lists, no links.',
  tools: [searchEvents, readPage],
  maxTurns: 6,
})
const foodScout = createSubagentTool({
  name: 'food_scout',
  description: 'Finds places to eat. Give it a full brief: it knows nothing else about the conversation.',
  provider,
  systemPrompt: 'You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.',
  tools: [searchPlaces, readReviews],
  maxTurns: 6,
})

// The parent only sees two tools. It never gets the listings, pages or reviews: just each scout's final text.
const agent = await createLocalAgent({
  provider,
  systemPrompt: 'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
  tools: [milongaScout, foodScout],
  maxTurns: 6,
})

const answer = await agent.run('I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.')
console.log(answer.content)
// Both scouts were asked for in one message, so the loop ran them in parallel.
// Cancelling the parent (abortSignal) cancels any scout still out.
```

**TypeScript**

```ts
// Subagents, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type Args = Record<string, string>
type ToolFn = (args: Args) => Promise<string>
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// 1. The loop from level 2, with its tools passed in. `messages` is born and dies inside each call.
async function runAgent(system: string, prompt: string, tools: Record<string, ToolFn>, schemas: object[], maxTurns = 6): Promise<string> {
  const messages: Message[] = [
    { role: 'system', content: system },
    { role: 'user', content: prompt },
  ]
  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: schemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    // Run every call of this message at once, and add the results in order.
    const calls = reply.tool_calls ?? []
    const outputs = await Promise.all(
      calls.map(async (call) => {
        try {
          const run = tools[call.function.name]
          return run ? await run(JSON.parse(call.function.arguments)) : `Unknown tool: ${call.function.name}`
        } catch (err) {
          return `Error: ${err instanceof Error ? err.message : err}`
        }
      }),
    )
    calls.forEach((call, i) => messages.push({ role: 'tool', tool_call_id: call.id, content: outputs[i]! }))
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// 2. The scouts' own tools: the heavy ones. Your code.
const scoutTools: Record<string, ToolFn> = {
  search_events: async ({ neighborhood, date }) => eventsApi.search(neighborhood, date),
  read_page: async ({ url }) => fetchText(url),
  search_places: async ({ near }) => placesApi.search(near),
  read_reviews: async ({ place }) => placesApi.reviews(place),
}
const pick = (...names: string[]) => Object.fromEntries(names.map((name) => [name, scoutTools[name]!]))

// 3. A subagent is a tool whose body is another runAgent call: new messages, fewer tools, its own prompt.
//    Only its final text comes back. Everything it read dies with its `messages`.
const parentTools: Record<string, ToolFn> = {
  milonga_scout: ({ task }) =>
    runAgent('You find milongas in Buenos Aires. Reply in 3 lines: name, address, times.', task, pick('search_events', 'read_page'), [/* their schemas */]),
  food_scout: ({ task }) =>
    runAgent('You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why.', task, pick('search_places', 'read_reviews'), [/* their schemas */]),
}
// Both take one string, `task`. The description tells the parent to write a full brief.
const parentSchemas = [/* milonga_scout(task), food_scout(task) */]

// 4. The parent: the same loop, and all it ever sees of the scouts is two short answers.
const answer = await runAgent(
  'You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.',
  'I’m staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.',
  parentTools,
  parentSchemas,
)
console.log(answer)
```

**Python**

```python
# Subagents, from scratch. Standard library only, no SDK.
import json
import urllib.request
from concurrent.futures import ThreadPoolExecutor

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

def post(path, payload):
    request = urllib.request.Request(
        f"{LLM['base_url']}{path}",
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

def call_tool(tools, call):
    try:
        return tools[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
    except Exception as err:
        return f"Error: {err}"

# 1. The loop from level 2, with its tools passed in. `messages` is born and dies inside each call.
def run_agent(system, prompt, tools, schemas, max_turns=6):
    messages = [{"role": "system", "content": system}, {"role": "user", "content": prompt}]

    for _ in range(max_turns):
        choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": schemas})["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        # Run every call of this message at once, and add the results in order.
        calls = reply.get("tool_calls", [])
        with ThreadPoolExecutor() as pool:
            outputs = list(pool.map(lambda call: call_tool(tools, call), calls))
        for call, output in zip(calls, outputs):
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# 2. The scouts' own tools: the heavy ones. Your code.
SCOUT_TOOLS = {
    "search_events": lambda neighborhood, date: events_api.search(neighborhood, date),
    "read_page": lambda url: fetch_text(url),
    "search_places": lambda near: places_api.search(near),
    "read_reviews": lambda place: places_api.reviews(place),
}

def pick(*names):
    return {name: SCOUT_TOOLS[name] for name in names}

# 3. A subagent is a tool whose body is another run_agent call: new messages, fewer tools, its own prompt.
#    Only its final text comes back. Everything it read dies with its `messages`.
def milonga_scout(task):
    system = "You find milongas in Buenos Aires. Reply in 3 lines: name, address, times."
    return run_agent(system, task, pick("search_events", "read_page"), [...])  # their schemas

def food_scout(task):
    system = "You find restaurants in Buenos Aires. Reply in 3 lines: name, address, why."
    return run_agent(system, task, pick("search_places", "read_reviews"), [...])  # their schemas

# Both take one string, `task`. The description tells the parent to write a full brief.
PARENT_TOOLS = {"milonga_scout": milonga_scout, "food_scout": food_scout}
PARENT_SCHEMAS = [...]  # milonga_scout(task), food_scout(task)

# 4. The parent: the same loop, and all it ever sees of the scouts is two short answers.
answer = run_agent(
    "You plan evenings out. Send the scouts out with a clear brief each, then put their answers together.",
    "I'm staying in San Telmo. Find me a milonga for Saturday night, and somewhere to eat nearby before it.",
    PARENT_TOOLS,
    PARENT_SCHEMAS,
)
print(answer)
```

## 注意事项

- **简报就是它知道的全部。**子智能体从没见过这段对话。“找我们刚才说的那个”对它毫无意义。在工具描述里要求父智能体写一份完整的简报：目标、约束条件，以及好答案应该是什么样子。
- **要求简短、固定的格式。**整件事的意义就在于得到一个小结果。像“用 3 行回答：名称、地址、时间”这样的 system prompt，能防止侦察员把它那堆东西又贴回给父智能体。
- **省下的是上下文，不是钱。**翻找照样在发生，只是发生在另一个智能体的请求里。总成本往往更高。当父智能体的专注值得这个价钱时才用子智能体，差事允许的话就给它们换个更便宜的模型。
- **只拆分彼此独立的部分。**两名侦察员能并排跑，是因为谁都不需要对方。如果第二件差事需要第一件的答案，就一个接一个地调用，或者干脆留在一个智能体里做。
- **限定它的工具，并限制嵌套深度。**只给每个子智能体它那件差事需要的工具；给它配属于它自己的子智能体之前，要三思。每多一层，调用次数就成倍增加，而深处的一个失败，传上来时只剩一行让人摸不着头脑的话。

## 相关模式

- [3 · 设计一个工具](https://harnesspatterns.dev/zh/patterns/designing-a-tool.md)
- [7 · 背包装满了](https://harnesspatterns.dev/zh/patterns/compaction.md)
- [8 · 按需加载的技能](https://harnesspatterns.dev/zh/patterns/skills.md)
- [10 · 每圈重新开始](https://harnesspatterns.dev/zh/patterns/fresh-laps.md)
- [11 · 可观测性与评估](https://harnesspatterns.dev/zh/patterns/observability.md)
- [14 · 安全与沙箱](https://harnesspatterns.dev/zh/patterns/security.md)
