Skip to content
astorlm
Language: English
← Map

Level 8

Prompt caching

Every turn, the loop sends the model the whole request again. Most of it is the same as last time, and the provider can remember it, as long as you don’t change how it starts.
1/35 Bandoneón folds:
  • user
  • assistant
  • tool_result
Game night. Every turn, the loop sends the Oracle the whole request again, so Astor plays it on Simon: one light per block, system prompt and tools first, then each message.

EventBus

The problem

A model remembers nothing between calls, so the loop resends everything on every turn: the system prompt, the tool list, and the whole history. In an agent that runs ten turns, the long instructions at the top travel ten times, and get paid ten times.

And they make it slower: before writing a single word, the model has to read the whole request again, from the first token.

The solution

Providers keep a short-lived memory of the requests they just read. When a new request starts exactly like a recent one, they reuse the part that matches and only process the rest. That part is billed at a fraction of the price (often a tenth, depending on the provider and the model), and the reply starts sooner.

The rule is in the word starts: the match is counted from the first token, and stops at the first one that differs. An agent fits it naturally, because each request is the last one plus a few messages at the end. All you have to do is not break it:

  • Something that changes, up top

    the time, a request id, the user’s name

    Keep the system prompt fixed. Put what changes in the latest message, at the end.

  • Tools that move

    a different order, a tool added mid-run

    Build the tool list once, in a fixed order, and send the same one every turn.

  • Rewriting the past

    editing, trimming or summarizing old messages

    Only append. When you do have to rewrite, everything after that point is paid again.

  • Switching models

    a cheaper model for one turn

    Each model has its own cache. A switch starts it from zero.

Some providers cache on their own once a request is long enough: OpenAI does it from 1,024 tokens. Others, like Anthropic, only cache where you mark it. Either way, the usage they return says how many input tokens came from the cache, so you can check.

The cast

Same cast as always, at game night.

Simon the request
Each round repeats the whole sequence and adds to its end, the way each turn resends the request.
The lights its blocks
Yellow for the system prompt and the tools, then one per message, in the colors of the history.
Astor the loop
Plays the whole sequence to the Oracle every turn, then runs the tools as usual.
The Oracle the model
Its bubble is the provider’s cache: the lights it already knows, from the first one on.
The scoreboard usage
Input tokens sent, read from cache, and paid, with cached ones at a tenth.
The living room the tools
The game shelf and the phone: game_shelf and order_pizza.

The code

With astorlm: the agent builds its system prompt once and resends it unchanged on every turn, with the same tools in the same order, and the history only grows at the end. Your part is to keep what changes out of systemPrompt. Each turn_end carries cacheReadTokens, which OpenAIProvider reads from the response.

From scratch: the loop from level 2, with its fixed start built once outside the loop, and a line that logs cached_tokens from each response’s usage.

import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const gameShelf = tool({
  name: 'game_shelf',
  description: 'List the board games on the living-room shelf.',
  schema: z.object({}),
  execute: async () => home.shelf(), // your code
})

const orderPizza = tool({
  name: 'order_pizza',
  description: 'Order pizza for delivery. Returns how long it will take.',
  schema: z.object({ size: z.enum(['medium', 'large']), count: z.number().int().min(1) }),
  execute: async ({ size, count }) => pizzeria.order(size, count), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  // The start of every request. astorlm builds it once and resends it unchanged every turn.
  // Nothing that changes goes here: no clock, no request id, no user name.
  systemPrompt: HOUSE_RULES, // a long, fixed text: the more of it, the more the cache saves
  tools: [gameShelf, orderPizza], // same tools, same order, every turn
  maxTurns: 10,
})

// OpenAI caches long prompts on its own (from 1,024 tokens), and says how much it reused.
agent.on('event', (event) => {
  if (event.type !== 'turn_end' || !event.usage) return
  const { inputTokens, cacheReadTokens = 0 } = event.usage
  console.log(`turn ${event.turn}: ${cacheReadTokens} of ${inputTokens} input tokens from cache`)
})

// Need the time? Put it at the end, in the message, where it only changes what comes after it.
const now = new Date().toLocaleTimeString()
await agent.run(`Game night for four: see what games we have, and order pizza. (It is ${now}.)`)

What to watch

  • The cache doesn’t last. It lives for a few minutes without use. An agent that waits an hour for a person, or a heartbeat that ticks every two, pays its first request again in full.
  • Compaction and caching pull in opposite directions. Shrinking old messages, the next level, rewrites the past, so every token after the first change is paid in full on the next turn. Compact rarely, and in big steps.
  • Short prompts don’t get cached. Below the provider’s minimum, there’s nothing to save. Caching pays off with long system prompts, many tools, or long histories: exactly what agents have.
  • Measure it. If cacheReadTokens stays at zero from the second turn on, something at the start is changing. Compare two requests side by side and find the first difference.
  • It saves on input only. The tokens the model writes cost the same. In agents, input is usually most of the bill, so it’s still the biggest saving on the table.