> Agent Harness Patterns 第 7 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/compaction · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 背包装满了

每一轮都要把整段历史记录重新发给模型。聊天持续得够久，它就装不下了。压缩会在那之前，先把最旧的部分缩小。

## 问题

模型一次能读的内容是有限的。这个上限就是它的**上下文窗口**，以 token 计（token 是词的片段，每个大约四个字符）：许多小型本地模型是 8,000，大型托管模型则有几十万。一次请求里的所有东西都必须装得下：system prompt、到目前为止的每条消息、每个工具结果，还要给回复留出空间。

模型在两次调用之间什么都不记得，所以循环每一轮都要发送整段历史记录。一个既查订单又搜目录的客服聊天，会堆积起成千上万个模型早已用过的结果 token。每一轮都比上一轮更慢、更贵。总有一天，请求会装不下。

这就是 Gulp，上下文溢出。这时提供商会用一个错误拒绝请求；更糟的是，有些服务器会悄悄截掉最旧的部分来凑合。而最旧的部分，正是顾客说明自己指的是哪个订单的地方。

## 解决方案

每次调用模型之前，循环都会检查请求有多大。一旦超过某条线（这条线定在真实上限之下，好给回复留出空间），它就缩小历史记录。常见的做法有三种：

- **截断旧的工具结果**
   把模型已经用过的大块结果，换成一行注释加一小段预览。
   几乎不花钱，而工具结果通常是历史记录里最重的部分。如果模型又需要细节，它会再调用一次工具。
- **丢弃旧的往来**
   删掉最早的请求，连同其后直到下一个请求为止的所有内容。
   同样不花钱，但模型会彻底忘掉那一段对话。要保留最开头的那个请求，因为它往往说明了整个聊天的主题。
- **让模型来总结**
   把旧的部分发给模型一次，用它的总结替换掉那些消息。
   能保留意思，但要多花一次调用，而且总结可能悄悄漏掉那个唯一要紧的数字。

不管选哪种，规则都一样：从最旧的消息开始，别动最近的几个请求，一装得下就停手。astorlm 按顺序做前两种：先截断旧的工具结果，只有不够时才丢弃消息。

## 角色

还是那群熟悉的角色，这次换到了下落方块井里。

- **井** (上下文窗口): 一次请求能装下的全部内容。如果方块堆到顶，请求就装不下了。
- **方块** (消息): 每条消息一个，大小取决于它的 token 数。井底是 system prompt：每次请求都带着它，而且永远不会被压缩。
- **红线** (阈值): 窗口的 80%。一旦越过，循环会在调用模型之前先压缩。
- **锤子** (压缩器): 把最旧的工具结果压成一行注释，上面的一切随之落定。
- **KEEP** (keepRecentTurns): 最近的两个请求，以及它们之后的一切。锤子永远不会碰它们。
- **SENT** (账单): 到目前为止所有调用累计发送的 token。对比一下锤子落下前后，它每一轮涨了多少。

在 EventBus 面板里，`compact` 这一行标记的是优化器在运行。astorlm 不会为此发出事件；它只会往你的 logger 里写一句 “Context optimized”。注意它出现的位置：在请求加入历史记录之后、`turn_start` 之前。

## 代码

**使用 astorlm：**压缩默认开启，大小根据提供商来定。传入 `contextOptimizer` 来设置真实的窗口、那条线，以及要保留多少个最近的请求。

**从零手写：**第 2 关的循环，加上一个跨请求保留的历史记录，并在每次调用模型之前调用 `compact()`。第 1 级截断旧的工具结果，第 2 级整段丢弃旧的往来。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const getOrder = tool({
  name: 'get_order',
  description: 'Everything about one order: items, shipping, invoice.',
  schema: z.object({ order: z.number().int() }),
  execute: async ({ order }) => loadOrder(order), // your code
})

const searchParts = tool({
  name: 'search_parts',
  description: 'Search the parts catalog, with stock and price for each match.',
  schema: z.object({ query: z.string() }),
  execute: async ({ query }) => searchCatalog(query), // your code
})

const createReturn = tool({
  name: 'create_return',
  description: 'Open a return for an order and ship a replacement part.',
  schema: z.object({ order: z.number().int(), part: z.string() }),
  execute: async ({ order, part }) => openReturn(order, part), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [getOrder, searchParts, createReturn],
  maxTurns: 10,
  // Compaction is on by default, sized from the provider. OpenAIProvider assumes a
  // 128,000-token window, so on a small local model, say how big it really is.
  contextOptimizer: {
    maxTokens: 8000,
    compressThreshold: 0.8, // compact once the request passes 80% of the window
    keepRecentTurns: 2, // never touch the last two requests, or anything after them
  },
  // There's no event for compaction: astorlm logs "Context optimized…" when it happens.
  logger: console,
})

// One agent, one history: every run() adds to it, and the optimizer checks it before each model call.
await agent.run('Hi! My order #4471 came with a bent front wheel. Can you help?')
await agent.run('Is that same wheel in stock?')
const last = await agent.run('Great. Open a return for my order and ship me the new wheel.')
console.log(last.content)
```

**TypeScript**

```ts
// Compaction, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

type ToolFn = (args: Record<string, string | number>) => Promise<string>
const tools: Record<string, ToolFn> = { get_order: getOrder, search_parts: searchParts, create_return: createReturn }
const toolSchemas = [/* one JSON Schema per tool */]

const SYSTEM = 'You are the support assistant of a bike shop.'
const WINDOW = { maxTokens: 8000, threshold: 0.8, keepRecentTurns: 2 }

// A rough count, about 4 characters per token. Good enough to decide when to compact.
function estimateTokens(messages: Message[]): number {
  const chars = messages.reduce((sum, m) => sum + JSON.stringify(m).length, SYSTEM.length)
  return Math.ceil(chars / 4)
}

// Where the protected part starts: the Nth user request from the end.
// With fewer requests than that, everything is recent and nothing can go.
function keepFrom(messages: Message[], keep: number): number {
  let seen = 0
  for (let i = messages.length - 1; i >= 0; i--) {
    if (messages[i]!.role === 'user' && ++seen === keep) return i
  }
  return 0
}

export function compact(messages: Message[]): Message[] {
  const limit = WINDOW.maxTokens * WINDOW.threshold
  if (estimateTokens(messages) <= limit) return messages
  const out = structuredClone(messages)

  // Level 1: shrink old tool results to a one-line note, oldest first. Stop as soon as it fits.
  for (let i = 0; i < keepFrom(out, WINDOW.keepRecentTurns); i++) {
    const m = out[i]!
    if (m.role !== 'tool' || m.content.startsWith('[Truncated')) continue
    m.content = `[Truncated to save context: ${m.content.length} chars. Preview: ${m.content.slice(0, 150)}…]`
    if (estimateTokens(out) <= limit) return out
  }

  // Level 2: drop the oldest exchanges whole, from one request up to the next.
  // Keep the very first request, and never cut between a tool call and its result.
  while (estimateTokens(out) > limit) {
    const next = out.findIndex((m, i) => i > 1 && m.role === 'user')
    if (next === -1 || next > keepFrom(out, WINDOW.keepRecentTurns)) break
    out.splice(1, next - 1)
  }
  return out
}

// The history lives across requests: that's what fills up.
const messages: Message[] = []

export async function ask(prompt: string, maxTurns = 10): Promise<string> {
  messages.push({ role: 'user', content: prompt })

  for (let turn = 1; turn <= maxTurns; turn++) {
    // Before every model call: does it still fit? The compacted history replaces the old one.
    messages.splice(0, messages.length, ...compact(messages))

    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: SYSTEM }, ...messages], tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      try {
        if (run) output = await run(JSON.parse(call.function.arguments))
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// The chat from the animation: three requests, one growing history.
await ask('Hi! My order #4471 came with a bent front wheel. Can you help?')
await ask('Is that same wheel in stock?')
console.log(await ask('Great. Open a return for my order and ship me the new wheel.'))
```

**Python**

```python
# Compaction, from scratch. Standard library only, no SDK.
import copy
import json
import math
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"get_order": get_order, "search_parts": search_parts, "create_return": create_return}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

SYSTEM = "You are the support assistant of a bike shop."
WINDOW = {"max_tokens": 8000, "threshold": 0.8, "keep_recent_turns": 2}

def estimate_tokens(messages):
    """A rough count, about 4 characters per token. Good enough to decide when to compact."""
    chars = len(SYSTEM) + sum(len(json.dumps(m)) for m in messages)
    return math.ceil(chars / 4)

def keep_from(messages, keep):
    """Where the protected part starts: the Nth user request from the end (0 if there are fewer)."""
    seen = 0
    for i in range(len(messages) - 1, -1, -1):
        if messages[i]["role"] == "user":
            seen += 1
            if seen == keep:
                return i
    return 0

def compact(messages):
    limit = WINDOW["max_tokens"] * WINDOW["threshold"]
    if estimate_tokens(messages) <= limit:
        return messages
    out = copy.deepcopy(messages)

    # Level 1: shrink old tool results to a one-line note, oldest first. Stop as soon as it fits.
    for i in range(keep_from(out, WINDOW["keep_recent_turns"])):
        m = out[i]
        if m["role"] != "tool" or m["content"].startswith("[Truncated"):
            continue
        m["content"] = f"[Truncated to save context: {len(m['content'])} chars. Preview: {m['content'][:150]}...]"
        if estimate_tokens(out) <= limit:
            return out

    # Level 2: drop the oldest exchanges whole, from one request up to the next.
    # Keep the very first request, and never cut between a tool call and its result.
    while estimate_tokens(out) > limit:
        following = [i for i, m in enumerate(out) if i > 1 and m["role"] == "user"]
        if not following or following[0] > keep_from(out, WINDOW["keep_recent_turns"]):
            break
        del out[1 : following[0]]
    return out

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({
            "model": LLM["model"],
            "messages": [{"role": "system", "content": SYSTEM}, *messages],
            "tools": TOOL_SCHEMAS,
        }).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

# The history lives across requests: that's what fills up.
messages = []

def ask(prompt, max_turns=10):
    messages.append({"role": "user", "content": prompt})

    for turn in range(1, max_turns + 1):
        # Before every model call: does it still fit? The compacted history replaces the old one.
        messages[:] = compact(messages)

        choice = chat(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            run = TOOLS.get(name)
            try:
                output = run(**json.loads(call["function"]["arguments"])) if run else f"Unknown tool: {name}"
            except Exception as err:
                output = f"Error: {err}"
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# The chat from the animation: three requests, one growing history.
ask("Hi! My order #4471 came with a bent front wheel. Can you help?")
ask("Is that same wheel in stock?")
print(ask("Great. Open a return for my order and ship me the new wheel."))
```

## 注意事项

- **告诉它真实的窗口大小。**astorlm 的 OpenAIProvider 默认假设 128,000 个 token。在一个 8,000 token 的本地模型上，优化器会一直等一条模型永远到不了的线。
- **永远不要把工具调用和它的结果拆开。**如果某个工具结果对应的调用被丢掉了，大多数 API 会拒绝整个请求。要按整段往来丢弃，从一个请求到下一个请求。
- **只有工具能再跑一次时，截断才安全。**如果某个结果没法再取一次，比如付款收据，就把要紧的部分留在答案里，或者存到历史记录之外。
- **告诉做总结的模型哪些东西必须保留。**订单号、零件号、已做的决定。一份读起来很顺的总结，照样可能丢掉下一轮唯一需要的那个事实。
- **压缩只计算你发送的内容。**按每个 token 四个字符来估算，足以判断何时动手。记得在线下给回复留够空间。

## 相关模式

- [2 · 智能体循环](https://harnesspatterns.dev/zh/patterns/agent-loop.md)
- [6 · 钩子](https://harnesspatterns.dev/zh/patterns/hooks.md)
- [8 · 按需加载的技能](https://harnesspatterns.dev/zh/patterns/skills.md)
- [9 · 记忆](https://harnesspatterns.dev/zh/patterns/memory.md)
- [15 · 子智能体](https://harnesspatterns.dev/zh/patterns/subagents.md)
