> Level 8 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/prompt-caching · All patterns: https://harnesspatterns.dev/llms.txt

# Prompt caching

Every turn, the loop sends the model the whole request again. Most of it is the same as last time, and the provider can remember it, as long as you don’t change how it starts.

## The problem

A model remembers nothing between calls, so the loop resends everything on every turn: the system prompt, the tool list, and the whole history. In an agent that runs ten turns, the long instructions at the top travel ten times, and get paid ten times.

And they make it slower: before writing a single word, the model has to read the whole request again, from the first token.

## The solution

Providers keep a short-lived memory of the requests they just read. When a new request starts exactly like a recent one, they reuse the part that matches and only process the rest. That part is billed at a fraction of the price (often a tenth, depending on the provider and the model), and the reply starts sooner.

The rule is in the word *starts*: the match is counted from the first token, and stops at the first one that differs. An agent fits it naturally, because each request is the last one plus a few messages at the end. All you have to do is not break it:

- **Something that changes, up top**
   the time, a request id, the user’s name
   Keep the system prompt fixed. Put what changes in the latest message, at the end.
- **Tools that move**
   a different order, a tool added mid-run
   Build the tool list once, in a fixed order, and send the same one every turn.
- **Rewriting the past**
   editing, trimming or summarizing old messages
   Only append. When you do have to rewrite, everything after that point is paid again.
- **Switching models**
   a cheaper model for one turn
   Each model has its own cache. A switch starts it from zero.

Some providers cache on their own once a request is long enough: OpenAI does it from 1,024 tokens. Others, like Anthropic, only cache where you mark it. Either way, the usage they return says how many input tokens came from the cache, so you can check.

## The cast

Same cast as always, at game night.

- **Simon** (the request): Each round repeats the whole sequence and adds to its end, the way each turn resends the request.
- **The lights** (its blocks): Yellow for the system prompt and the tools, then one per message, in the colors of the history.
- **Astor** (the loop): Plays the whole sequence to the Oracle every turn, then runs the tools as usual.
- **The Oracle** (the model): Its bubble is the provider’s cache: the lights it already knows, from the first one on.
- **The scoreboard** (usage): Input tokens sent, read from cache, and paid, with cached ones at a tenth.
- **The living room** (the tools): The game shelf and the phone: `game_shelf` and `order_pizza`.

## The code

**With astorlm:** the agent builds its system prompt once and resends it unchanged on every turn, with the same tools in the same order, and the history only grows at the end. Your part is to keep what changes out of `systemPrompt`. Each `turn_end` carries `cacheReadTokens`, which `OpenAIProvider` reads from the response.

**From scratch:** the loop from level 2, with its fixed start built once outside the loop, and a line that logs `cached_tokens` from each response’s usage.

**With astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const gameShelf = tool({
  name: 'game_shelf',
  description: 'List the board games on the living-room shelf.',
  schema: z.object({}),
  execute: async () => home.shelf(), // your code
})

const orderPizza = tool({
  name: 'order_pizza',
  description: 'Order pizza for delivery. Returns how long it will take.',
  schema: z.object({ size: z.enum(['medium', 'large']), count: z.number().int().min(1) }),
  execute: async ({ size, count }) => pizzeria.order(size, count), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  // The start of every request. astorlm builds it once and resends it unchanged every turn.
  // Nothing that changes goes here: no clock, no request id, no user name.
  systemPrompt: HOUSE_RULES, // a long, fixed text: the more of it, the more the cache saves
  tools: [gameShelf, orderPizza], // same tools, same order, every turn
  maxTurns: 10,
})

// OpenAI caches long prompts on its own (from 1,024 tokens), and says how much it reused.
agent.on('event', (event) => {
  if (event.type !== 'turn_end' || !event.usage) return
  const { inputTokens, cacheReadTokens = 0 } = event.usage
  console.log(`turn ${event.turn}: ${cacheReadTokens} of ${inputTokens} input tokens from cache`)
})

// Need the time? Put it at the end, in the message, where it only changes what comes after it.
const now = new Date().toLocaleTimeString()
await agent.run(`Game night for four: see what games we have, and order pizza. (It is ${now}.)`)
```

**TypeScript**

```ts
// Prompt caching from scratch. Plain fetch, no SDK.
// There is nothing to build: the provider caches. Your job is to not break it.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

const tools: Record<string, (args: Record<string, string>) => Promise<string>> = { game_shelf: gameShelf, order_pizza: orderPizza }

// 1. The fixed start, built once: same text and same tools, in the same order, every request.
const SYSTEM: Message = { role: 'system', content: HOUSE_RULES } // ✗ never `It is ${new Date()}` up here
const TOOL_SCHEMAS = Object.freeze([/* one JSON Schema per tool, always in this order */])

export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  // 2. The history only grows at the end. Editing an old message breaks the cache from there on.
  const messages: Message[] = [SYSTEM, { role: 'user', content: prompt }]

  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: TOOL_SCHEMAS }),
    })
    const { choices, usage } = await res.json()

    // 3. Check it's working: how much of the input the provider read from its cache.
    const cached = usage?.prompt_tokens_details?.cached_tokens ?? 0
    console.log(`turn ${turn}: ${cached} of ${usage?.prompt_tokens} input tokens from cache`)

    const [choice] = choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls' || reply.role !== 'assistant') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const output = await tools[call.function.name]!(JSON.parse(call.function.arguments || '{}'))
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}
```

**Python**

```python
# Prompt caching from scratch. Standard library only, no SDK.
# There is nothing to build: the provider caches. Your job is to not break it.
import json
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"game_shelf": game_shelf, "order_pizza": order_pizza}

# 1. The fixed start, built once: same text and same tools, in the same order, every request.
SYSTEM = {"role": "system", "content": HOUSE_RULES}  # never f"It is {datetime.now()}" up here
TOOL_SCHEMAS = (...)  # one JSON Schema per tool, always in this order

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({"model": LLM["model"], "messages": messages, "tools": list(TOOL_SCHEMAS)}).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

def run_agent(prompt, max_turns=10):
    # 2. The history only grows at the end. Editing an old message breaks the cache from there on.
    messages = [SYSTEM, {"role": "user", "content": prompt}]

    for turn in range(1, max_turns + 1):
        data = chat(messages)

        # 3. Check it's working: how much of the input the provider read from its cache.
        usage = data.get("usage") or {}
        cached = (usage.get("prompt_tokens_details") or {}).get("cached_tokens", 0)
        print(f"turn {turn}: {cached} of {usage.get('prompt_tokens')} input tokens from cache")

        choice = data["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"] or "{}"))
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")
```

## What to watch

- **The cache doesn’t last.** It lives for a few minutes without use. An agent that waits an hour for a person, or a heartbeat that ticks every two, pays its first request again in full.
- **Compaction and caching pull in opposite directions.** Shrinking old messages, the next level, rewrites the past, so every token after the first change is paid in full on the next turn. Compact rarely, and in big steps.
- **Short prompts don’t get cached.** Below the provider’s minimum, there’s nothing to save. Caching pays off with long system prompts, many tools, or long histories: exactly what agents have.
- **Measure it.** If `cacheReadTokens` stays at zero from the second turn on, something at the start is changing. Compare two requests side by side and find the first difference.
- **It saves on input only.** The tokens the model writes cost the same. In agents, input is usually most of the bill, so it’s still the biggest saving on the table.

## Related patterns

- [0 · Your toolkit](https://harnesspatterns.dev/patterns/your-toolkit.md)
- [2 · Agent Loop](https://harnesspatterns.dev/patterns/agent-loop.md)
- [9 · The backpack fills up](https://harnesspatterns.dev/patterns/compaction.md)
- [15 · Observability and evals](https://harnesspatterns.dev/patterns/observability.md)
