Level 10
Fresh laps
- user
- assistant
- tool_result
EventBus
The problem
Some jobs don’t fit in one run: migrate 300 files, translate a whole catalog, fix every failing test in a repo. Each step adds a tool call and a tool result to the history, and the loop resends all of it every turn.
That’s Muddle, the endless session. By the afternoon it’s carrying every step since the morning: the requests are huge, the old results bury the new ones, and the model starts redoing work it already did or skipping work it only planned. Nothing crashes. The quality just drains away.
Compaction (level 7) slows Muddle down. It doesn’t stop it: a long enough job ends up summarizing its own summaries.
The solution
Don’t keep one agent alive for the whole job. Run it in laps. Each lap, your code starts a new agent with an empty history and the same goal. It does one slice of the work, writes down where things stand, and ends. Then your code checks the work itself and, if it isn’t done, starts the next lap.
-
One long session
Keep the same agent and the same history for the whole job.
Every turn resends everything since the start. The requests get heavier, the model gets worse at finding what matters in them, and past the window it breaks.
-
Compact as you go
Same session, but shrink the old messages when the history gets close to the limit (level 7).
Buys time, not a fix. Every compaction loses detail, and a long enough job compacts its own summaries.
-
Fresh laps
Split the job into laps. Each lap is a new agent with an empty history. What it needs to know, it reads from files.
Every lap starts small and clean. The price: each lap spends a turn or two finding its bearings, and the files have to say everything that matters.
The trick is that nothing important lives in the history. The work is on disk (the bridge), and so is a short
note saying how far it got (PROGRESS.md). A new agent doesn’t need to remember the last lap. It
only needs to read.
The pattern is often called the Ralph loop, after a shell one-liner that fed a coding agent the same prompt over and over. Coding agents use it for long refactors, with the git tree and a TODO file as the state.
The cast
Same cast as always, in a canyon this time.
- The hatch your code
-
Starts a new agent every lap (
createIterationAgent) and gets its answer back. It’s the loop around the loop. - A lap’s Astor one agent run
- The agent loop from level 2, with its own bandoneón. It starts empty and floats away when the lap ends.
- The bridge the work
- What the tools changed on disk. No lap ever throws it away.
- The sign PROGRESS.md
- A short note from each lap to the next: what’s done, what comes next.
- DONE? isDone
- Your check, between laps. It measures the bridge, not what the model says about it.
- LAP 3/5 maxIterations
- The fuse. If the job never checks out, the loop stops anyway.
Watch the two bars at the top. This lap is what each request really weighs, and it starts over every lap. 1 session is what the same requests would weigh if one agent had done all three laps: it never goes down.
The code
With astorlm: runGoalLoop takes a factory that returns a new agent, your
isDone check and a maxIterations fuse. The tools write to files, so every lap finds the
work where the last one left it.
From scratch: The loop from level 2, called inside a for. The history is a local
variable of each call, so every lap starts empty for free.
import { OpenAIProvider, createLocalAgent, runGoalLoop, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// The state lives on disk, not in any history: the bridge, and a progress note.
const GAP = 36
const bridgeLength = (): number => (existsSync('bridge.json') ? JSON.parse(readFileSync('bridge.json', 'utf8')).length : 0)
const readProgress = tool({
name: 'read_progress',
description: 'Read PROGRESS.md: what earlier laps built, and where to start.',
schema: z.object({}),
execute: async () => (existsSync('PROGRESS.md') ? readFileSync('PROGRESS.md', 'utf8') : 'Nothing built yet.'),
})
const layBricks = tool({
name: 'lay_bricks',
description: 'Lay up to 12 bricks of the bridge, starting at brick number "from".',
schema: z.object({ from: z.number().int().min(1), count: z.number().int().min(1).max(12) }),
execute: async ({ from, count }) => {
const to = Math.min(from + count - 1, GAP)
writeFileSync('bridge.json', JSON.stringify({ length: Math.max(bridgeLength(), to) }))
return `Laid bricks ${from}-${to}. The bridge is ${bridgeLength()} bricks long.`
},
})
const writeProgress = tool({
name: 'write_progress',
description: 'Overwrite PROGRESS.md with where the bridge stands now, for whoever comes next.',
schema: z.object({ text: z.string() }),
execute: async ({ text }) => {
writeFileSync('PROGRESS.md', `# Progress\n${text}\n`)
return 'Saved PROGRESS.md.'
},
})
const result = await runGoalLoop({
goal: 'Build the bridge to the exit: 36 bricks. Read PROGRESS.md first, lay at most 12 bricks, then update PROGRESS.md.',
// A NEW agent every lap: empty history, fresh context window. Same tools, same folder.
createIterationAgent: () =>
createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
tools: [readProgress, layBricks, writeProgress],
maxTurns: 8,
}),
// Your code decides when the job is done, by checking the work itself. Not the model's word.
isDone: () => bridgeLength() >= GAP,
onIteration: ({ iteration, lastText }) => console.log(`lap ${iteration}: ${lastText}`),
maxIterations: 5, // the fuse: a goal that never checks out can't run forever
})
console.log(result) // { iterations: 3, done: true, stopReason: 'done', lastText: '…' }
// Fresh laps, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. The state lives on disk: the bridge, and a progress note for the next lap.
const GAP = 36
const bridgeLength = (): number => (existsSync('bridge.json') ? JSON.parse(readFileSync('bridge.json', 'utf8')).length : 0)
type ToolFn = (args: Record<string, string | number>) => string
const tools: Record<string, ToolFn> = {
read_progress: () => (existsSync('PROGRESS.md') ? readFileSync('PROGRESS.md', 'utf8') : 'Nothing built yet.'),
lay_bricks: ({ from, count }) => {
const to = Math.min(Number(from) + Math.min(Number(count), 12) - 1, GAP)
writeFileSync('bridge.json', JSON.stringify({ length: Math.max(bridgeLength(), to) }))
return `Laid bricks ${from}-${to}. The bridge is ${bridgeLength()} bricks long.`
},
write_progress: ({ text }) => {
writeFileSync('PROGRESS.md', `# Progress\n${text}\n`)
return 'Saved PROGRESS.md.'
},
}
const toolSchemas = [/* one JSON Schema per tool: read_progress(), lay_bricks(from, count), write_progress(text) */]
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 2. The loop from level 2, unchanged. `messages` is born and dies inside each call.
async function runAgent(prompt: string, maxTurns = 8): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// 3. The goal loop: a fresh run per lap, then YOUR check of the work on disk.
const GOAL = 'Build the bridge to the exit: 36 bricks. Read PROGRESS.md first, lay at most 12 bricks, then update PROGRESS.md.'
const MAX_LAPS = 5 // the fuse
for (let lap = 1; lap <= MAX_LAPS; lap++) {
console.log(`lap ${lap}:`, await runAgent(GOAL))
if (bridgeLength() >= GAP) {
console.log(`Done after ${lap} laps.`)
break
}
if (lap === MAX_LAPS) throw new Error(`Bridge unfinished after ${MAX_LAPS} laps: ${bridgeLength()}/${GAP}`)
}
# Fresh laps, from scratch. Standard library only, no SDK.
import json
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# 1. The state lives on disk: the bridge, and a progress note for the next lap.
GAP = 36
BRIDGE = Path("bridge.json")
PROGRESS = Path("PROGRESS.md")
def bridge_length():
return json.loads(BRIDGE.read_text())["length"] if BRIDGE.exists() else 0
def read_progress():
return PROGRESS.read_text() if PROGRESS.exists() else "Nothing built yet."
def lay_bricks(start, count):
end = min(start + min(count, 12) - 1, GAP)
BRIDGE.write_text(json.dumps({"length": max(bridge_length(), end)}))
return f"Laid bricks {start}-{end}. The bridge is {bridge_length()} bricks long."
def write_progress(text):
PROGRESS.write_text(f"# Progress\n{text}\n")
return "Saved PROGRESS.md."
TOOLS = {
"read_progress": read_progress,
"lay_bricks": lambda **args: lay_bricks(args["from"], args["count"]), # "from" is a Python keyword
"write_progress": write_progress,
}
TOOL_SCHEMAS = [...] # one JSON Schema per tool: read_progress(), lay_bricks(from, count), write_progress(text)
# 2. The loop from level 2, unchanged. `messages` is born and dies inside each call.
def run_agent(prompt, max_turns=8):
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
try:
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# 3. The goal loop: a fresh run per lap, then YOUR check of the work on disk.
GOAL = "Build the bridge to the exit: 36 bricks. Read PROGRESS.md first, lay at most 12 bricks, then update PROGRESS.md."
MAX_LAPS = 5 # the fuse
for lap in range(1, MAX_LAPS + 1):
print(f"lap {lap}:", run_agent(GOAL))
if bridge_length() >= GAP:
print(f"Done after {lap} laps.")
break
else:
raise RuntimeError(f"Bridge unfinished after {MAX_LAPS} laps: {bridge_length()}/{GAP}")
What to watch
-
Check the work, not the answer. “Done!” from the model proves nothing.
isDoneshould look at the result itself: run the tests, count the rows, measure the bridge. Keep it cheap and deterministic, because it runs after every lap. -
Always set the fuse. A check that can never pass, or an agent that keeps undoing its own work,
loops until your bill stops it.
maxIterations, and a look at why it ran out. - The progress file is the only handover. Whatever it leaves out, the next lap doesn’t know. Tell the agent exactly what to write there: what’s done, what’s next, what it tried that failed.
- Make each step safe to repeat. A lap can die halfway, after the work but before the note. The next lap will do that slice again, so doing it twice must not break anything.
- Keep the slices small. A lap should fit in one short run. If a single slice already needs compaction, the slices are too big.