Level 8
Prompt caching
- user
- assistant
- tool_result
EventBus
The problem
A model remembers nothing between calls, so the loop resends everything on every turn: the system prompt, the tool list, and the whole history. In an agent that runs ten turns, the long instructions at the top travel ten times, and get paid ten times.
And they make it slower: before writing a single word, the model has to read the whole request again, from the first token.
The solution
Providers keep a short-lived memory of the requests they just read. When a new request starts exactly like a recent one, they reuse the part that matches and only process the rest. That part is billed at a fraction of the price (often a tenth, depending on the provider and the model), and the reply starts sooner.
The rule is in the word starts: the match is counted from the first token, and stops at the first one that differs. An agent fits it naturally, because each request is the last one plus a few messages at the end. All you have to do is not break it:
-
Something that changes, up top
the time, a request id, the user’s name
Keep the system prompt fixed. Put what changes in the latest message, at the end.
-
Tools that move
a different order, a tool added mid-run
Build the tool list once, in a fixed order, and send the same one every turn.
-
Rewriting the past
editing, trimming or summarizing old messages
Only append. When you do have to rewrite, everything after that point is paid again.
-
Switching models
a cheaper model for one turn
Each model has its own cache. A switch starts it from zero.
Some providers cache on their own once a request is long enough: OpenAI does it from 1,024 tokens. Others, like Anthropic, only cache where you mark it. Either way, the usage they return says how many input tokens came from the cache, so you can check.
The cast
Same cast as always, at game night.
- Simon the request
- Each round repeats the whole sequence and adds to its end, the way each turn resends the request.
- The lights its blocks
- Yellow for the system prompt and the tools, then one per message, in the colors of the history.
- Astor the loop
- Plays the whole sequence to the Oracle every turn, then runs the tools as usual.
- The Oracle the model
- Its bubble is the provider’s cache: the lights it already knows, from the first one on.
- The scoreboard usage
- Input tokens sent, read from cache, and paid, with cached ones at a tenth.
- The living room the tools
- The game shelf and the phone:
game_shelfandorder_pizza.
The code
With astorlm: the agent builds its system prompt once and resends it unchanged on every turn, with
the same tools in the same order, and the history only grows at the end. Your part is to keep what changes out of
systemPrompt. Each turn_end carries cacheReadTokens, which
OpenAIProvider reads from the response.
From scratch: the loop from level 2, with its fixed start built once outside the loop, and a line
that logs cached_tokens from each response’s usage.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
const gameShelf = tool({
name: 'game_shelf',
description: 'List the board games on the living-room shelf.',
schema: z.object({}),
execute: async () => home.shelf(), // your code
})
const orderPizza = tool({
name: 'order_pizza',
description: 'Order pizza for delivery. Returns how long it will take.',
schema: z.object({ size: z.enum(['medium', 'large']), count: z.number().int().min(1) }),
execute: async ({ size, count }) => pizzeria.order(size, count), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
// The start of every request. astorlm builds it once and resends it unchanged every turn.
// Nothing that changes goes here: no clock, no request id, no user name.
systemPrompt: HOUSE_RULES, // a long, fixed text: the more of it, the more the cache saves
tools: [gameShelf, orderPizza], // same tools, same order, every turn
maxTurns: 10,
})
// OpenAI caches long prompts on its own (from 1,024 tokens), and says how much it reused.
agent.on('event', (event) => {
if (event.type !== 'turn_end' || !event.usage) return
const { inputTokens, cacheReadTokens = 0 } = event.usage
console.log(`turn ${event.turn}: ${cacheReadTokens} of ${inputTokens} input tokens from cache`)
})
// Need the time? Put it at the end, in the message, where it only changes what comes after it.
const now = new Date().toLocaleTimeString()
await agent.run(`Game night for four: see what games we have, and order pizza. (It is ${now}.)`)
// Prompt caching from scratch. Plain fetch, no SDK.
// There is nothing to build: the provider caches. Your job is to not break it.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
const tools: Record<string, (args: Record<string, string>) => Promise<string>> = { game_shelf: gameShelf, order_pizza: orderPizza }
// 1. The fixed start, built once: same text and same tools, in the same order, every request.
const SYSTEM: Message = { role: 'system', content: HOUSE_RULES } // ✗ never `It is ${new Date()}` up here
const TOOL_SCHEMAS = Object.freeze([/* one JSON Schema per tool, always in this order */])
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
// 2. The history only grows at the end. Editing an old message breaks the cache from there on.
const messages: Message[] = [SYSTEM, { role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: TOOL_SCHEMAS }),
})
const { choices, usage } = await res.json()
// 3. Check it's working: how much of the input the provider read from its cache.
const cached = usage?.prompt_tokens_details?.cached_tokens ?? 0
console.log(`turn ${turn}: ${cached} of ${usage?.prompt_tokens} input tokens from cache`)
const [choice] = choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls' || reply.role !== 'assistant') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const output = await tools[call.function.name]!(JSON.parse(call.function.arguments || '{}'))
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
# Prompt caching from scratch. Standard library only, no SDK.
# There is nothing to build: the provider caches. Your job is to not break it.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"game_shelf": game_shelf, "order_pizza": order_pizza}
# 1. The fixed start, built once: same text and same tools, in the same order, every request.
SYSTEM = {"role": "system", "content": HOUSE_RULES} # never f"It is {datetime.now()}" up here
TOOL_SCHEMAS = (...) # one JSON Schema per tool, always in this order
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": list(TOOL_SCHEMAS)}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
def run_agent(prompt, max_turns=10):
# 2. The history only grows at the end. Editing an old message breaks the cache from there on.
messages = [SYSTEM, {"role": "user", "content": prompt}]
for turn in range(1, max_turns + 1):
data = chat(messages)
# 3. Check it's working: how much of the input the provider read from its cache.
usage = data.get("usage") or {}
cached = (usage.get("prompt_tokens_details") or {}).get("cached_tokens", 0)
print(f"turn {turn}: {cached} of {usage.get('prompt_tokens')} input tokens from cache")
choice = data["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"] or "{}"))
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
What to watch
- The cache doesn’t last. It lives for a few minutes without use. An agent that waits an hour for a person, or a heartbeat that ticks every two, pays its first request again in full.
- Compaction and caching pull in opposite directions. Shrinking old messages, the next level, rewrites the past, so every token after the first change is paid in full on the next turn. Compact rarely, and in big steps.
- Short prompts don’t get cached. Below the provider’s minimum, there’s nothing to save. Caching pays off with long system prompts, many tools, or long histories: exactly what agents have.
-
Measure it. If
cacheReadTokensstays at zero from the second turn on, something at the start is changing. Compare two requests side by side and find the first difference. - It saves on input only. The tokens the model writes cost the same. In agents, input is usually most of the bill, so it’s still the biggest saving on the table.