> Level 7 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/compaction · All patterns: https://harnesspatterns.dev/llms.txt

# The backpack fills up

Every turn sends the whole history to the model again. Keep a chat going long enough and it stops fitting. Compaction shrinks the oldest parts before that happens.

## The problem

A model can only read so much at once. That limit is its **context window**, counted in tokens (pieces of words, about four characters each): 8,000 for many small local models, a few hundred thousand for the big hosted ones. Everything in one request has to fit in it: the system prompt, every message so far, every tool result, and room for the reply.

The model remembers nothing between calls, so the loop sends the whole history every turn. A support chat that looks up an order and searches a catalog piles up thousands of tokens of results the model already used. Each turn is slower and costs more than the last. And one day the request doesn’t fit.

That’s Gulp, the context overflow. Then the provider rejects the request with an error, or, worse, some servers quietly cut off the oldest part to make it fit. The oldest part is where the customer said which order they meant.

## The solution

Before every call to the model, the loop checks how big the request is. Past a line, set below the real limit so the reply still has room, it shrinks the history. There are three common ways to do it:

- **Truncate old tool results**
   Replace a big result the model has already used with a one-line note and a short preview.
   Almost free, and tool results are usually the heaviest part of the history. If the model needs the details again, it calls the tool again.
- **Drop old exchanges**
   Remove the oldest requests with everything that came after them, up to the next request.
   Free too, but the model forgets that part of the conversation completely. Keep the very first request, since it often says what the whole chat is about.
- **Summarize with the model**
   Send the old part to the model once, and put its summary in place of those messages.
   It keeps the meaning, but costs an extra call, and a summary can quietly leave out the one number that mattered.

Whichever you pick, the same rules apply: start with the oldest messages, leave the last few requests alone, and stop as soon as it fits. astorlm does the first two, in that order. It truncates old tool results first, and only drops messages if that wasn’t enough.

## The cast

Same cast as always, in a falling-blocks well this time.

- **The well** (context window): Everything one request can carry. If the stack reaches the top, the request doesn’t fit.
- **The blocks** (messages): One per message, sized by its tokens. The floor is the system prompt: it goes with every request and never gets compacted.
- **The red line** (threshold): 80% of the window. Past it, the loop compacts before it calls the model.
- **The hammer** (the compactor): Shrinks the oldest tool result to a one-line note, and everything above it settles.
- **KEEP** (keepRecentTurns): The last two requests and everything after them. The hammer never touches them.
- **SENT** (the bill): Tokens sent so far, over every call. Watch how much it grows per turn before and after the hammer.

In the EventBus panel, the `compact` line marks the optimizer running. astorlm doesn’t emit an event for it; it writes “Context optimized” to your logger. Notice where it falls: after the request joins the history, before `turn_start`.

## The code

**With astorlm:** Compaction is on by default, sized from the provider. Pass `contextOptimizer` to set the real window, the line, and how many recent requests to keep.

**From scratch:** The loop from level 2, with a history that lives across requests and a `compact()` call before every model call. Level 1 truncates old tool results, level 2 drops old exchanges whole.

**With astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

const getOrder = tool({
  name: 'get_order',
  description: 'Everything about one order: items, shipping, invoice.',
  schema: z.object({ order: z.number().int() }),
  execute: async ({ order }) => loadOrder(order), // your code
})

const searchParts = tool({
  name: 'search_parts',
  description: 'Search the parts catalog, with stock and price for each match.',
  schema: z.object({ query: z.string() }),
  execute: async ({ query }) => searchCatalog(query), // your code
})

const createReturn = tool({
  name: 'create_return',
  description: 'Open a return for an order and ship a replacement part.',
  schema: z.object({ order: z.number().int(), part: z.string() }),
  execute: async ({ order, part }) => openReturn(order, part), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [getOrder, searchParts, createReturn],
  maxTurns: 10,
  // Compaction is on by default, sized from the provider. OpenAIProvider assumes a
  // 128,000-token window, so on a small local model, say how big it really is.
  contextOptimizer: {
    maxTokens: 8000,
    compressThreshold: 0.8, // compact once the request passes 80% of the window
    keepRecentTurns: 2, // never touch the last two requests, or anything after them
  },
  // There's no event for compaction: astorlm logs "Context optimized…" when it happens.
  logger: console,
})

// One agent, one history: every run() adds to it, and the optimizer checks it before each model call.
await agent.run('Hi! My order #4471 came with a bent front wheel. Can you help?')
await agent.run('Is that same wheel in stock?')
const last = await agent.run('Great. Open a return for my order and ship me the new wheel.')
console.log(last.content)
```

**TypeScript**

```ts
// Compaction, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

type ToolFn = (args: Record<string, string | number>) => Promise<string>
const tools: Record<string, ToolFn> = { get_order: getOrder, search_parts: searchParts, create_return: createReturn }
const toolSchemas = [/* one JSON Schema per tool */]

const SYSTEM = 'You are the support assistant of a bike shop.'
const WINDOW = { maxTokens: 8000, threshold: 0.8, keepRecentTurns: 2 }

// A rough count, about 4 characters per token. Good enough to decide when to compact.
function estimateTokens(messages: Message[]): number {
  const chars = messages.reduce((sum, m) => sum + JSON.stringify(m).length, SYSTEM.length)
  return Math.ceil(chars / 4)
}

// Where the protected part starts: the Nth user request from the end.
// With fewer requests than that, everything is recent and nothing can go.
function keepFrom(messages: Message[], keep: number): number {
  let seen = 0
  for (let i = messages.length - 1; i >= 0; i--) {
    if (messages[i]!.role === 'user' && ++seen === keep) return i
  }
  return 0
}

export function compact(messages: Message[]): Message[] {
  const limit = WINDOW.maxTokens * WINDOW.threshold
  if (estimateTokens(messages) <= limit) return messages
  const out = structuredClone(messages)

  // Level 1: shrink old tool results to a one-line note, oldest first. Stop as soon as it fits.
  for (let i = 0; i < keepFrom(out, WINDOW.keepRecentTurns); i++) {
    const m = out[i]!
    if (m.role !== 'tool' || m.content.startsWith('[Truncated')) continue
    m.content = `[Truncated to save context: ${m.content.length} chars. Preview: ${m.content.slice(0, 150)}…]`
    if (estimateTokens(out) <= limit) return out
  }

  // Level 2: drop the oldest exchanges whole, from one request up to the next.
  // Keep the very first request, and never cut between a tool call and its result.
  while (estimateTokens(out) > limit) {
    const next = out.findIndex((m, i) => i > 1 && m.role === 'user')
    if (next === -1 || next > keepFrom(out, WINDOW.keepRecentTurns)) break
    out.splice(1, next - 1)
  }
  return out
}

// The history lives across requests: that's what fills up.
const messages: Message[] = []

export async function ask(prompt: string, maxTurns = 10): Promise<string> {
  messages.push({ role: 'user', content: prompt })

  for (let turn = 1; turn <= maxTurns; turn++) {
    // Before every model call: does it still fit? The compacted history replaces the old one.
    messages.splice(0, messages.length, ...compact(messages))

    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: SYSTEM }, ...messages], tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      try {
        if (run) output = await run(JSON.parse(call.function.arguments))
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// The chat from the animation: three requests, one growing history.
await ask('Hi! My order #4471 came with a bent front wheel. Can you help?')
await ask('Is that same wheel in stock?')
console.log(await ask('Great. Open a return for my order and ship me the new wheel.'))
```

**Python**

```python
# Compaction, from scratch. Standard library only, no SDK.
import copy
import json
import math
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

TOOLS = {"get_order": get_order, "search_parts": search_parts, "create_return": create_return}
TOOL_SCHEMAS = [...]  # one JSON Schema per tool

SYSTEM = "You are the support assistant of a bike shop."
WINDOW = {"max_tokens": 8000, "threshold": 0.8, "keep_recent_turns": 2}

def estimate_tokens(messages):
    """A rough count, about 4 characters per token. Good enough to decide when to compact."""
    chars = len(SYSTEM) + sum(len(json.dumps(m)) for m in messages)
    return math.ceil(chars / 4)

def keep_from(messages, keep):
    """Where the protected part starts: the Nth user request from the end (0 if there are fewer)."""
    seen = 0
    for i in range(len(messages) - 1, -1, -1):
        if messages[i]["role"] == "user":
            seen += 1
            if seen == keep:
                return i
    return 0

def compact(messages):
    limit = WINDOW["max_tokens"] * WINDOW["threshold"]
    if estimate_tokens(messages) <= limit:
        return messages
    out = copy.deepcopy(messages)

    # Level 1: shrink old tool results to a one-line note, oldest first. Stop as soon as it fits.
    for i in range(keep_from(out, WINDOW["keep_recent_turns"])):
        m = out[i]
        if m["role"] != "tool" or m["content"].startswith("[Truncated"):
            continue
        m["content"] = f"[Truncated to save context: {len(m['content'])} chars. Preview: {m['content'][:150]}...]"
        if estimate_tokens(out) <= limit:
            return out

    # Level 2: drop the oldest exchanges whole, from one request up to the next.
    # Keep the very first request, and never cut between a tool call and its result.
    while estimate_tokens(out) > limit:
        following = [i for i, m in enumerate(out) if i > 1 and m["role"] == "user"]
        if not following or following[0] > keep_from(out, WINDOW["keep_recent_turns"]):
            break
        del out[1 : following[0]]
    return out

def chat(messages):
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({
            "model": LLM["model"],
            "messages": [{"role": "system", "content": SYSTEM}, *messages],
            "tools": TOOL_SCHEMAS,
        }).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)["choices"][0]

# The history lives across requests: that's what fills up.
messages = []

def ask(prompt, max_turns=10):
    messages.append({"role": "user", "content": prompt})

    for turn in range(1, max_turns + 1):
        # Before every model call: does it still fit? The compacted history replaces the old one.
        messages[:] = compact(messages)

        choice = chat(messages)
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            run = TOOLS.get(name)
            try:
                output = run(**json.loads(call["function"]["arguments"])) if run else f"Unknown tool: {name}"
            except Exception as err:
                output = f"Error: {err}"
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# The chat from the animation: three requests, one growing history.
ask("Hi! My order #4471 came with a bent front wheel. Can you help?")
ask("Is that same wheel in stock?")
print(ask("Great. Open a return for my order and ship me the new wheel."))
```

## What to watch

- **Tell it the real window.** astorlm’s OpenAIProvider assumes 128,000 tokens. On an 8,000-token local model, the optimizer would wait for a line the model never reaches.
- **Never split a tool call from its result.** A tool result whose call was dropped makes most APIs reject the whole request. Drop whole exchanges, from one request up to the next.
- **Truncating is only safe if the tool can run again.** If a result can’t be fetched twice, like a payment receipt, keep the part that matters in the answer, or save it outside the history.
- **Tell the summarizer what must survive.** Order numbers, part numbers, decisions. A summary that reads well can still lose the one fact the next turn needs.
- **Compaction only counts what you send.** Estimating four characters per token is fine for deciding when to act. Leave enough room under the line for the reply.

## Related patterns

- [2 · Agent Loop](https://harnesspatterns.dev/patterns/agent-loop.md)
- [6 · Hooks](https://harnesspatterns.dev/patterns/hooks.md)
- [8 · On-demand skills](https://harnesspatterns.dev/patterns/skills.md)
- [9 · Memory](https://harnesspatterns.dev/patterns/memory.md)
- [15 · Subagents](https://harnesspatterns.dev/patterns/subagents.md)
