第 8 关
提示词缓存
- user
- assistant
- tool_result
EventBus
问题
模型在两次调用之间什么都不记得,所以循环每一轮都要重发一切:system prompt、工具列表和整段历史。一个跑十轮的智能体里,最上面那段长长的指令要传十次,也要付十次钱。
而且还会更慢:在写出第一个字之前,模型得从第一个 token 开始,把整个请求重新读一遍。
解决方案
服务商会对刚读过的请求保留一份短期记忆。当一个新请求的开头和最近的某个请求一模一样时,它们会复用对得上的那部分,只处理剩下的。那部分只按零头收费(常常是十分之一,取决于服务商和模型),回复也会更早开始。
规则就在“开头”二字:匹配从第一个 token 算起,遇到第一个不同的 token 就停。智能体天然契合这一点,因为每个请求都是上一个请求再在末尾加几条消息。你要做的只是别把它打破:
-
最上面放了会变的东西
时间、请求 id、用户名
让 system prompt 保持固定。会变的东西放进最新的消息里,放在最后。
-
工具挪来挪去
换了顺序,或者中途加了一个工具
工具列表只构建一次,顺序固定,每一轮都发同一份。
-
改写过去
编辑、截断或总结旧消息
只在末尾追加。真要改写时,那个位置之后的所有内容都要重新付费。
-
换模型
某一轮换个更便宜的模型
每个模型都有自己的缓存。一换就从零开始。
有些服务商在请求够长时会自动缓存:OpenAI 从 1,024 个 token 起就会这样做。另一些,比如 Anthropic,只在你标记的地方缓存。无论哪种,它们返回的用量都会告诉你有多少输入 token 来自缓存,你可以自己核对。
角色
还是那群熟悉的角色,这次是游戏之夜。
- Simon 请求
- 每一回合都重复整段序列,再在末尾加点东西,就像每一轮都重发请求一样。
- 灯 请求的各个块
- 黄色代表 system prompt 和工具,之后每条消息一盏,颜色和历史记录一致。
- Astor 循环
- 每一轮都把整段序列弹给神谕者听,然后照常运行工具。
- 神谕者 模型
- 它的气泡就是服务商的缓存:从第一盏起,它已经认识的那些灯。
- 记分板 用量
- 发送的、从缓存读取的和实际付费的输入 token,缓存部分按十分之一计。
- 客厅 工具
- 游戏架和电话:
game_shelf和order_pizza。
代码
使用 astorlm:智能体只构建一次 system prompt,之后每一轮原样重发,工具也一样、顺序也一样,历史只在末尾增长。你要做的是把会变的东西挡在 systemPrompt 外面。每个 turn_end 都带着 cacheReadTokens,由 OpenAIProvider 从响应里读出来。
从零手写:第 2 关的循环,固定的开头在循环外只构建一次,再加一行代码,从每次响应的用量里记录 cached_tokens。
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const gameShelf = tool({
name: 'game_shelf',
description: 'List the board games on the living-room shelf.',
schema: z.object({}),
execute: async () => home.shelf(), // your code
})
const orderPizza = tool({
name: 'order_pizza',
description: 'Order pizza for delivery. Returns how long it will take.',
schema: z.object({ size: z.enum(['medium', 'large']), count: z.number().int().min(1) }),
execute: async ({ size, count }) => pizzeria.order(size, count), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
// The start of every request. astorlm builds it once and resends it unchanged every turn.
// Nothing that changes goes here: no clock, no request id, no user name.
systemPrompt: HOUSE_RULES, // a long, fixed text: the more of it, the more the cache saves
tools: [gameShelf, orderPizza], // same tools, same order, every turn
maxTurns: 10,
})
// OpenAI caches long prompts on its own (from 1,024 tokens), and says how much it reused.
agent.on('event', (event) => {
if (event.type !== 'turn_end' || !event.usage) return
const { inputTokens, cacheReadTokens = 0 } = event.usage
console.log(`turn ${event.turn}: ${cacheReadTokens} of ${inputTokens} input tokens from cache`)
})
// Need the time? Put it at the end, in the message, where it only changes what comes after it.
const now = new Date().toLocaleTimeString()
await agent.run(`Game night for four: see what games we have, and order pizza. (It is ${now}.)`)
// Prompt caching from scratch. Plain fetch, no SDK.
// There is nothing to build: the provider caches. Your job is to not break it.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
const tools: Record<string, (args: Record<string, string>) => Promise<string>> = { game_shelf: gameShelf, order_pizza: orderPizza }
// 1. The fixed start, built once: same text and same tools, in the same order, every request.
const SYSTEM: Message = { role: 'system', content: HOUSE_RULES } // ✗ never `It is ${new Date()}` up here
const TOOL_SCHEMAS = Object.freeze([/* one JSON Schema per tool, always in this order */])
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
// 2. The history only grows at the end. Editing an old message breaks the cache from there on.
const messages: Message[] = [SYSTEM, { role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: TOOL_SCHEMAS }),
})
const { choices, usage } = await res.json()
// 3. Check it's working: how much of the input the provider read from its cache.
const cached = usage?.prompt_tokens_details?.cached_tokens ?? 0
console.log(`turn ${turn}: ${cached} of ${usage?.prompt_tokens} input tokens from cache`)
const [choice] = choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls' || reply.role !== 'assistant') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const output = await tools[call.function.name]!(JSON.parse(call.function.arguments || '{}'))
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
# Prompt caching from scratch. Standard library only, no SDK.
# There is nothing to build: the provider caches. Your job is to not break it.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"game_shelf": game_shelf, "order_pizza": order_pizza}
# 1. The fixed start, built once: same text and same tools, in the same order, every request.
SYSTEM = {"role": "system", "content": HOUSE_RULES} # never f"It is {datetime.now()}" up here
TOOL_SCHEMAS = (...) # one JSON Schema per tool, always in this order
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": list(TOOL_SCHEMAS)}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
def run_agent(prompt, max_turns=10):
# 2. The history only grows at the end. Editing an old message breaks the cache from there on.
messages = [SYSTEM, {"role": "user", "content": prompt}]
for turn in range(1, max_turns + 1):
data = chat(messages)
# 3. Check it's working: how much of the input the provider read from its cache.
usage = data.get("usage") or {}
cached = (usage.get("prompt_tokens_details") or {}).get("cached_tokens", 0)
print(f"turn {turn}: {cached} of {usage.get('prompt_tokens')} input tokens from cache")
choice = data["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"] or "{}"))
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
注意事项
- 缓存不会长久。闲置几分钟它就没了。一个要等人一个小时的智能体,或者每两小时跳一次的心跳,第一个请求又得全价付费。
- 压缩和缓存往相反的方向使劲。缩减旧消息(也就是下一关)会改写过去,于是第一处改动之后的每个 token,在下一轮都要全价付费。压缩要少做,一次压多一点。
- 短提示词不会被缓存。低于服务商的最低长度,就没什么可省的。缓存的价值在于长 system prompt、大量工具或长历史:恰好是智能体都有的东西。
-
去测量。如果从第二轮起
cacheReadTokens一直是零,说明开头有东西在变。把两个请求并排比较,找出第一处不同。 - 它只省输入。模型写出来的 token 价格不变。在智能体里,输入通常占账单的大头,所以这依然是能拿到的最大一笔节省。