> Agent Harness Patterns 第 8 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/prompt-caching · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 提示词缓存

每一轮，循环都会把整个请求重新发给模型。其中大部分和上一次一样，服务商能记住它，只要你别改动它的开头。

## 问题

模型在两次调用之间什么都不记得，所以循环每一轮都要重发一切：system prompt、工具列表和整段历史。一个跑十轮的智能体里，最上面那段长长的指令要传十次，也要付十次钱。

而且还会更慢：在写出第一个字之前，模型得从第一个 token 开始，把整个请求重新读一遍。

## 解决方案

服务商会对刚读过的请求保留一份短期记忆。当一个新请求的开头和最近的某个请求一模一样时，它们会复用对得上的那部分，只处理剩下的。那部分只按零头收费（常常是十分之一，取决于服务商和模型），回复也会更早开始。

规则就在“开头”二字：匹配从第一个 token 算起，遇到第一个不同的 token 就停。智能体天然契合这一点，因为每个请求都是上一个请求再在末尾加几条消息。你要做的只是别把它打破：

- **最上面放了会变的东西**
   时间、请求 id、用户名
   让 system prompt 保持固定。会变的东西放进最新的消息里，放在最后。
- **工具挪来挪去**
   换了顺序，或者中途加了一个工具
   工具列表只构建一次，顺序固定，每一轮都发同一份。
- **改写过去**
   编辑、截断或总结旧消息
   只在末尾追加。真要改写时，那个位置之后的所有内容都要重新付费。
- **换模型**
   某一轮换个更便宜的模型
   每个模型都有自己的缓存。一换就从零开始。

有些服务商在请求够长时会自动缓存：OpenAI 从 1,024 个 token 起就会这样做。另一些，比如 Anthropic，只在你标记的地方缓存。无论哪种，它们返回的用量都会告诉你有多少输入 token 来自缓存，你可以自己核对。

## 角色

还是那群熟悉的角色，这次是游戏之夜。

- **Simon** (请求): 每一回合都重复整段序列，再在末尾加点东西，就像每一轮都重发请求一样。
- **灯** (请求的各个块): 黄色代表 system prompt 和工具，之后每条消息一盏，颜色和历史记录一致。
- **Astor** (循环): 每一轮都把整段序列弹给神谕者听，然后照常运行工具。
- **神谕者** (模型): 它的气泡就是服务商的缓存：从第一盏起，它已经认识的那些灯。
- **记分板** (用量): 发送的、从缓存读取的和实际付费的输入 token，缓存部分按十分之一计。
- **客厅** (工具): 游戏架和电话：`game_shelf` 和 `order_pizza`。

## 代码

**使用 astorlm：**智能体只构建一次 system prompt，之后每一轮原样重发，工具也一样、顺序也一样，历史只在末尾增长。你要做的是把会变的东西挡在 `systemPrompt` 外面。每个 `turn_end` 都带着 `cacheReadTokens`，由 `OpenAIProvider` 从响应里读出来。

**从零手写：**第 2 关的循环，固定的开头在循环外只构建一次，再加一行代码，从每次响应的用量里记录 `cached_tokens`。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const gameShelf = tool({
  name: 'game_shelf',
  description: 'List the board games on the living-room shelf.',
  schema: z.object({}),
  execute: async () => home.shelf(), // your code
})

const orderPizza = tool({
  name: 'order_pizza',
  description: 'Order pizza for delivery. Returns how long it will take.',
  schema: z.object({ size: z.enum(['medium', 'large']), count: z.number().int().min(1) }),
  execute: async ({ size, count }) => pizzeria.order(size, count), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  // The start of every request. astorlm builds it once and resends it unchanged every turn.
  // Nothing that changes goes here: no clock, no request id, no user name.
  systemPrompt: HOUSE_RULES, // a long, fixed text: the more of it, the more the cache saves
  tools: [gameShelf, orderPizza], // same tools, same order, every turn
  maxTurns: 10,
})

// OpenAI caches long prompts on its own (from 1,024 tokens), and says how much it reused.
agent.on('event', (event) => {
  if (event.type !== 'turn_end' || !event.usage) return
  const { inputTokens, cacheReadTokens = 0 } = event.usage
  console.log(`turn ${event.turn}: ${cacheReadTokens} of ${inputTokens} input tokens from cache`)
})

// Need the time? Put it at the end, in the message, where it only changes what comes after it.
const now = new Date().toLocaleTimeString()
await agent.run(`Game night for four: see what games we have, and order pizza. (It is ${now}.)`)
```

**TypeScript**

```ts
// Prompt caching from scratch. Plain fetch, no SDK.
// There is nothing to build: the provider caches. Your job is to not break it.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

const tools: Record<string, (args: Record<string, string>) => Promise<string>> = { game_shelf: gameShelf, order_pizza: orderPizza }

// 1. The fixed start, built once: same text and same tools, in the same order, every request.
const SYSTEM: Message = { role: 'system', content: HOUSE_RULES } // ✗ never `It is ${new Date()}` up here
const TOOL_SCHEMAS = Object.freeze([/* one JSON Schema per tool, always in this order */])

export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  // 2. The history only grows at the end. Editing an old message breaks the cache from there on.
  const messages: Message[] = [SYSTEM, { role: 'user', content: prompt }]

  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: TOOL_SCHEMAS }),
    })
    const { choices, usage } = await res.json()

    // 3. Check it's working: how much of the input the provider read from its cache.
    const cached = usage?.prompt_tokens_details?.cached_tokens ?? 0
    console.log(`turn ${turn}: ${cached} of ${usage?.prompt_tokens} input tokens from cache`)

    const [choice] = choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls' || reply.role !== 'assistant') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const output = await tools[call.function.name]!(JSON.parse(call.function.arguments || '{}'))
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
```

**Python**

```python
# Prompt caching from scratch. Standard library only, no SDK.
# There is nothing to build: the provider caches. Your job is to not break it.
import json
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"game_shelf": game_shelf, "order_pizza": order_pizza}

# 1. The fixed start, built once: same text and same tools, in the same order, every request.
SYSTEM = {"role": "system", "content": HOUSE_RULES}  # never f"It is {datetime.now()}" up here
TOOL_SCHEMAS = (...)  # one JSON Schema per tool, always in this order

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": list(TOOL_SCHEMAS)}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

def run_agent(prompt, max_turns=10):
    # 2. The history only grows at the end. Editing an old message breaks the cache from there on.
    messages = [SYSTEM, {"role": "user", "content": prompt}]

    for turn in range(1, max_turns + 1):
        data = chat(messages)

        # 3. Check it's working: how much of the input the provider read from its cache.
        usage = data.get("usage") or {}
        cached = (usage.get("prompt_tokens_details") or {}).get("cached_tokens", 0)
        print(f"turn {turn}: {cached} of {usage.get('prompt_tokens')} input tokens from cache")

        choice = data["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"] or "{}"))
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")
```

## 注意事项

- **缓存不会长久。**闲置几分钟它就没了。一个要等人一个小时的智能体，或者每两小时跳一次的心跳，第一个请求又得全价付费。
- **压缩和缓存往相反的方向使劲。**缩减旧消息（也就是下一关）会改写过去，于是第一处改动之后的每个 token，在下一轮都要全价付费。压缩要少做，一次压多一点。
- **短提示词不会被缓存。**低于服务商的最低长度，就没什么可省的。缓存的价值在于长 system prompt、大量工具或长历史：恰好是智能体都有的东西。
- **去测量。**如果从第二轮起 `cacheReadTokens` 一直是零，说明开头有东西在变。把两个请求并排比较，找出第一处不同。
- **它只省输入。**模型写出来的 token 价格不变。在智能体里，输入通常占账单的大头，所以这依然是能拿到的最大一笔节省。

## 相关模式

- [0 · 你的工具箱](https://harnesspatterns.dev/zh/patterns/your-toolkit.md)
- [2 · 智能体循环](https://harnesspatterns.dev/zh/patterns/agent-loop.md)
- [9 · 背包装满了](https://harnesspatterns.dev/zh/patterns/compaction.md)
- [15 · 可观测性与评估](https://harnesspatterns.dev/zh/patterns/observability.md)
