Skip to content
astorlm
← Map

Level 7

The backpack fills up

Every turn sends the whole history to the model again. Keep a chat going long enough and it stops fitting. Compaction shrinks the oldest parts before that happens.
1/30 Messages:
  • user
  • assistant
  • tool_result
A bike shop’s support assistant, on a model that reads 8,000 tokens at most. The well is that window. Every message drops in as a block, and the system prompt is the floor. The red line is at 80%.

EventBus

The problem

A model can only read so much at once. That limit is its context window, counted in tokens (pieces of words, about four characters each): 8,000 for many small local models, a few hundred thousand for the big hosted ones. Everything in one request has to fit in it: the system prompt, every message so far, every tool result, and room for the reply.

The model remembers nothing between calls, so the loop sends the whole history every turn. A support chat that looks up an order and searches a catalog piles up thousands of tokens of results the model already used. Each turn is slower and costs more than the last. And one day the request doesn’t fit.

That’s Gulp, the context overflow. Then the provider rejects the request with an error, or, worse, some servers quietly cut off the oldest part to make it fit. The oldest part is where the customer said which order they meant.

The solution

Before every call to the model, the loop checks how big the request is. Past a line, set below the real limit so the reply still has room, it shrinks the history. There are three common ways to do it:

  • Truncate old tool results

    Replace a big result the model has already used with a one-line note and a short preview.

    Almost free, and tool results are usually the heaviest part of the history. If the model needs the details again, it calls the tool again.

  • Drop old exchanges

    Remove the oldest requests with everything that came after them, up to the next request.

    Free too, but the model forgets that part of the conversation completely. Keep the very first request, since it often says what the whole chat is about.

  • Summarize with the model

    Send the old part to the model once, and put its summary in place of those messages.

    It keeps the meaning, but costs an extra call, and a summary can quietly leave out the one number that mattered.

Whichever you pick, the same rules apply: start with the oldest messages, leave the last few requests alone, and stop as soon as it fits. astorlm does the first two, in that order. It truncates old tool results first, and only drops messages if that wasn’t enough.

The cast

Same cast as always, in a falling-blocks well this time.

The well context window
Everything one request can carry. If the stack reaches the top, the request doesn’t fit.
The blocks messages
One per message, sized by its tokens. The floor is the system prompt: it goes with every request and never gets compacted.
The red line threshold
80% of the window. Past it, the loop compacts before it calls the model.
The hammer the compactor
Shrinks the oldest tool result to a one-line note, and everything above it settles.
KEEP keepRecentTurns
The last two requests and everything after them. The hammer never touches them.
SENT the bill
Tokens sent so far, over every call. Watch how much it grows per turn before and after the hammer.

In the EventBus panel, the compact line marks the optimizer running. astorlm doesn’t emit an event for it; it writes “Context optimized” to your logger. Notice where it falls: after the request joins the history, before turn_start.

The code

With astorlm: Compaction is on by default, sized from the provider. Pass contextOptimizer to set the real window, the line, and how many recent requests to keep.

From scratch: The loop from level 2, with a history that lives across requests and a compact() call before every model call. Level 1 truncates old tool results, level 2 drops old exchanges whole.

import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const getOrder = tool({
  name: 'get_order',
  description: 'Everything about one order: items, shipping, invoice.',
  schema: z.object({ order: z.number().int() }),
  execute: async ({ order }) => loadOrder(order), // your code
})

const searchParts = tool({
  name: 'search_parts',
  description: 'Search the parts catalog, with stock and price for each match.',
  schema: z.object({ query: z.string() }),
  execute: async ({ query }) => searchCatalog(query), // your code
})

const createReturn = tool({
  name: 'create_return',
  description: 'Open a return for an order and ship a replacement part.',
  schema: z.object({ order: z.number().int(), part: z.string() }),
  execute: async ({ order, part }) => openReturn(order, part), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [getOrder, searchParts, createReturn],
  maxTurns: 10,
  // Compaction is on by default, sized from the provider. OpenAIProvider assumes a
  // 128,000-token window, so on a small local model, say how big it really is.
  contextOptimizer: {
    maxTokens: 8000,
    compressThreshold: 0.8, // compact once the request passes 80% of the window
    keepRecentTurns: 2, // never touch the last two requests, or anything after them
  },
  // There's no event for compaction: astorlm logs "Context optimized…" when it happens.
  logger: console,
})

// One agent, one history: every run() adds to it, and the optimizer checks it before each model call.
await agent.run('Hi! My order #4471 came with a bent front wheel. Can you help?')
await agent.run('Is that same wheel in stock?')
const last = await agent.run('Great. Open a return for my order and ship me the new wheel.')
console.log(last.content)

What to watch

  • Tell it the real window. astorlm’s OpenAIProvider assumes 128,000 tokens. On an 8,000-token local model, the optimizer would wait for a line the model never reaches.
  • Never split a tool call from its result. A tool result whose call was dropped makes most APIs reject the whole request. Drop whole exchanges, from one request up to the next.
  • Truncating is only safe if the tool can run again. If a result can’t be fetched twice, like a payment receipt, keep the part that matters in the answer, or save it outside the history.
  • Tell the summarizer what must survive. Order numbers, part numbers, decisions. A summary that reads well can still lose the one fact the next turn needs.
  • Compaction only counts what you send. Estimating four characters per token is fine for deciding when to act. Leave enough room under the line for the reply.