第 7 关
背包装满了
- user
- assistant
- tool_result
EventBus
问题
模型一次能读的内容是有限的。这个上限就是它的上下文窗口,以 token 计(token 是词的片段,每个大约四个字符):许多小型本地模型是 8,000,大型托管模型则有几十万。一次请求里的所有东西都必须装得下:system prompt、到目前为止的每条消息、每个工具结果,还要给回复留出空间。
模型在两次调用之间什么都不记得,所以循环每一轮都要发送整段历史记录。一个既查订单又搜目录的客服聊天,会堆积起成千上万个模型早已用过的结果 token。每一轮都比上一轮更慢、更贵。总有一天,请求会装不下。
这就是 Gulp,上下文溢出。这时提供商会用一个错误拒绝请求;更糟的是,有些服务器会悄悄截掉最旧的部分来凑合。而最旧的部分,正是顾客说明自己指的是哪个订单的地方。
解决方案
每次调用模型之前,循环都会检查请求有多大。一旦超过某条线(这条线定在真实上限之下,好给回复留出空间),它就缩小历史记录。常见的做法有三种:
-
截断旧的工具结果
把模型已经用过的大块结果,换成一行注释加一小段预览。
几乎不花钱,而工具结果通常是历史记录里最重的部分。如果模型又需要细节,它会再调用一次工具。
-
丢弃旧的往来
删掉最早的请求,连同其后直到下一个请求为止的所有内容。
同样不花钱,但模型会彻底忘掉那一段对话。要保留最开头的那个请求,因为它往往说明了整个聊天的主题。
-
让模型来总结
把旧的部分发给模型一次,用它的总结替换掉那些消息。
能保留意思,但要多花一次调用,而且总结可能悄悄漏掉那个唯一要紧的数字。
不管选哪种,规则都一样:从最旧的消息开始,别动最近的几个请求,一装得下就停手。astorlm 按顺序做前两种:先截断旧的工具结果,只有不够时才丢弃消息。
角色
还是那群熟悉的角色,这次换到了下落方块井里。
- 井 上下文窗口
- 一次请求能装下的全部内容。如果方块堆到顶,请求就装不下了。
- 方块 消息
- 每条消息一个,大小取决于它的 token 数。井底是 system prompt:每次请求都带着它,而且永远不会被压缩。
- 红线 阈值
- 窗口的 80%。一旦越过,循环会在调用模型之前先压缩。
- 锤子 压缩器
- 把最旧的工具结果压成一行注释,上面的一切随之落定。
- KEEP keepRecentTurns
- 最近的两个请求,以及它们之后的一切。锤子永远不会碰它们。
- SENT 账单
- 到目前为止所有调用累计发送的 token。对比一下锤子落下前后,它每一轮涨了多少。
在 EventBus 面板里,compact 这一行标记的是优化器在运行。astorlm 不会为此发出事件;它只会往你的 logger 里写一句 “Context optimized”。注意它出现的位置:在请求加入历史记录之后、turn_start 之前。
代码
使用 astorlm:压缩默认开启,大小根据提供商来定。传入 contextOptimizer 来设置真实的窗口、那条线,以及要保留多少个最近的请求。
从零手写:第 2 关的循环,加上一个跨请求保留的历史记录,并在每次调用模型之前调用 compact()。第 1 级截断旧的工具结果,第 2 级整段丢弃旧的往来。
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const getOrder = tool({
name: 'get_order',
description: 'Everything about one order: items, shipping, invoice.',
schema: z.object({ order: z.number().int() }),
execute: async ({ order }) => loadOrder(order), // your code
})
const searchParts = tool({
name: 'search_parts',
description: 'Search the parts catalog, with stock and price for each match.',
schema: z.object({ query: z.string() }),
execute: async ({ query }) => searchCatalog(query), // your code
})
const createReturn = tool({
name: 'create_return',
description: 'Open a return for an order and ship a replacement part.',
schema: z.object({ order: z.number().int(), part: z.string() }),
execute: async ({ order, part }) => openReturn(order, part), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [getOrder, searchParts, createReturn],
maxTurns: 10,
// Compaction is on by default, sized from the provider. OpenAIProvider assumes a
// 128,000-token window, so on a small local model, say how big it really is.
contextOptimizer: {
maxTokens: 8000,
compressThreshold: 0.8, // compact once the request passes 80% of the window
keepRecentTurns: 2, // never touch the last two requests, or anything after them
},
// There's no event for compaction: astorlm logs "Context optimized…" when it happens.
logger: console,
})
// One agent, one history: every run() adds to it, and the optimizer checks it before each model call.
await agent.run('Hi! My order #4471 came with a bent front wheel. Can you help?')
await agent.run('Is that same wheel in stock?')
const last = await agent.run('Great. Open a return for my order and ship me the new wheel.')
console.log(last.content)
// Compaction, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, string | number>) => Promise<string>
const tools: Record<string, ToolFn> = { get_order: getOrder, search_parts: searchParts, create_return: createReturn }
const toolSchemas = [/* one JSON Schema per tool */]
const SYSTEM = 'You are the support assistant of a bike shop.'
const WINDOW = { maxTokens: 8000, threshold: 0.8, keepRecentTurns: 2 }
// A rough count, about 4 characters per token. Good enough to decide when to compact.
function estimateTokens(messages: Message[]): number {
const chars = messages.reduce((sum, m) => sum + JSON.stringify(m).length, SYSTEM.length)
return Math.ceil(chars / 4)
}
// Where the protected part starts: the Nth user request from the end.
// With fewer requests than that, everything is recent and nothing can go.
function keepFrom(messages: Message[], keep: number): number {
let seen = 0
for (let i = messages.length - 1; i >= 0; i--) {
if (messages[i]!.role === 'user' && ++seen === keep) return i
}
return 0
}
export function compact(messages: Message[]): Message[] {
const limit = WINDOW.maxTokens * WINDOW.threshold
if (estimateTokens(messages) <= limit) return messages
const out = structuredClone(messages)
// Level 1: shrink old tool results to a one-line note, oldest first. Stop as soon as it fits.
for (let i = 0; i < keepFrom(out, WINDOW.keepRecentTurns); i++) {
const m = out[i]!
if (m.role !== 'tool' || m.content.startsWith('[Truncated')) continue
m.content = `[Truncated to save context: ${m.content.length} chars. Preview: ${m.content.slice(0, 150)}…]`
if (estimateTokens(out) <= limit) return out
}
// Level 2: drop the oldest exchanges whole, from one request up to the next.
// Keep the very first request, and never cut between a tool call and its result.
while (estimateTokens(out) > limit) {
const next = out.findIndex((m, i) => i > 1 && m.role === 'user')
if (next === -1 || next > keepFrom(out, WINDOW.keepRecentTurns)) break
out.splice(1, next - 1)
}
return out
}
// The history lives across requests: that's what fills up.
const messages: Message[] = []
export async function ask(prompt: string, maxTurns = 10): Promise<string> {
messages.push({ role: 'user', content: prompt })
for (let turn = 1; turn <= maxTurns; turn++) {
// Before every model call: does it still fit? The compacted history replaces the old one.
messages.splice(0, messages.length, ...compact(messages))
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: SYSTEM }, ...messages], tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// The chat from the animation: three requests, one growing history.
await ask('Hi! My order #4471 came with a bent front wheel. Can you help?')
await ask('Is that same wheel in stock?')
console.log(await ask('Great. Open a return for my order and ship me the new wheel.'))
# Compaction, from scratch. Standard library only, no SDK.
import copy
import json
import math
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"get_order": get_order, "search_parts": search_parts, "create_return": create_return}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
SYSTEM = "You are the support assistant of a bike shop."
WINDOW = {"max_tokens": 8000, "threshold": 0.8, "keep_recent_turns": 2}
def estimate_tokens(messages):
"""A rough count, about 4 characters per token. Good enough to decide when to compact."""
chars = len(SYSTEM) + sum(len(json.dumps(m)) for m in messages)
return math.ceil(chars / 4)
def keep_from(messages, keep):
"""Where the protected part starts: the Nth user request from the end (0 if there are fewer)."""
seen = 0
for i in range(len(messages) - 1, -1, -1):
if messages[i]["role"] == "user":
seen += 1
if seen == keep:
return i
return 0
def compact(messages):
limit = WINDOW["max_tokens"] * WINDOW["threshold"]
if estimate_tokens(messages) <= limit:
return messages
out = copy.deepcopy(messages)
# Level 1: shrink old tool results to a one-line note, oldest first. Stop as soon as it fits.
for i in range(keep_from(out, WINDOW["keep_recent_turns"])):
m = out[i]
if m["role"] != "tool" or m["content"].startswith("[Truncated"):
continue
m["content"] = f"[Truncated to save context: {len(m['content'])} chars. Preview: {m['content'][:150]}...]"
if estimate_tokens(out) <= limit:
return out
# Level 2: drop the oldest exchanges whole, from one request up to the next.
# Keep the very first request, and never cut between a tool call and its result.
while estimate_tokens(out) > limit:
following = [i for i, m in enumerate(out) if i > 1 and m["role"] == "user"]
if not following or following[0] > keep_from(out, WINDOW["keep_recent_turns"]):
break
del out[1 : following[0]]
return out
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({
"model": LLM["model"],
"messages": [{"role": "system", "content": SYSTEM}, *messages],
"tools": TOOL_SCHEMAS,
}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
# The history lives across requests: that's what fills up.
messages = []
def ask(prompt, max_turns=10):
messages.append({"role": "user", "content": prompt})
for turn in range(1, max_turns + 1):
# Before every model call: does it still fit? The compacted history replaces the old one.
messages[:] = compact(messages)
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
run = TOOLS.get(name)
try:
output = run(**json.loads(call["function"]["arguments"])) if run else f"Unknown tool: {name}"
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# The chat from the animation: three requests, one growing history.
ask("Hi! My order #4471 came with a bent front wheel. Can you help?")
ask("Is that same wheel in stock?")
print(ask("Great. Open a return for my order and ship me the new wheel."))
注意事项
- 告诉它真实的窗口大小。astorlm 的 OpenAIProvider 默认假设 128,000 个 token。在一个 8,000 token 的本地模型上,优化器会一直等一条模型永远到不了的线。
- 永远不要把工具调用和它的结果拆开。如果某个工具结果对应的调用被丢掉了,大多数 API 会拒绝整个请求。要按整段往来丢弃,从一个请求到下一个请求。
- 只有工具能再跑一次时,截断才安全。如果某个结果没法再取一次,比如付款收据,就把要紧的部分留在答案里,或者存到历史记录之外。
- 告诉做总结的模型哪些东西必须保留。订单号、零件号、已做的决定。一份读起来很顺的总结,照样可能丢掉下一轮唯一需要的那个事实。
- 压缩只计算你发送的内容。按每个 token 四个字符来估算,足以判断何时动手。记得在线下给回复留够空间。