Level 12
Plan and reflect
- user
- assistant
- tool_result
EventBus
The problem
Give an agent a job with several parts and it starts with whatever is in front of it. Halfway through, the early steps are far back in the history, and it forgets one. At the end it answers “done!” with total confidence, because nothing forces it to look.
That’s Scatterbrain. Two failures in one: no plan, so steps get lost; no check, so a step that went wrong gets reported as done. On Tango Street the tool said plainly that the paper landed in the bushes. The model read it and ticked the box anyway.
The solution
Plan first. Before the first real action, the agent writes the work down as a list of tasks, and
ticks each one as it goes. The trick that makes it work: the whole plan is added to every request, so the model
always sees what’s done and what’s left, however long the history gets. In astorlm that’s
pattern: 'PLAN_EXECUTE': two tools, add_plan_item and update_plan_item, and the
plan in the system prompt on every turn.
Reflect before accepting. A ticked box is only what the model says. When run()
returns, your code reviews the result before accepting it. If something is missing, the finding goes back as a new
message on the same session: the agent keeps its history and its plan, and fixes only what’s wrong. Cap the
number of rounds. There are three ways to review, from strongest to weakest:
-
Check the world
Your code looks at the result itself: the porches, the database rows, the test suite, the file on disk.
The best reviewer when you can have it: cheap, exact, and it can’t be talked into anything. It needs the work to be checkable by code.
-
A critic model
A second call reads the task, the answer and a checklist, and lists what’s wrong or missing.
For work no code can check: a summary, an email, a plan. It costs a call, it can miss things too, and it needs concrete criteria, not “is this good?”.
-
Ask the agent
The system prompt tells the agent to re-read its work before it answers.
Free and sometimes enough. But it’s the same model grading itself, with the same blind spots: it ticked #14 once already.
This is the same idea as an eval from level 11, used at runtime: an eval grades runs after the fact to improve the agent; a review grades this run before the user ever sees it.
The cast
Same cast as always, on a paper round this time.
- The kiosk the model
- The Oracle, behind the counter. It decides every step, and never rides.
- The route sheet the plan
- The tasks and their boxes. It flashes gold on every turn: it goes into every request.
- Astor’s bike the loop
- Rides each tool call out and back, with the history in his bandoneón.
- A throw deliver
- An ordinary tool. It says where the paper landed.
- The editor your code
-
Hands over the job, and reviews the porches before accepting the answer. Its magnifier is
review().
The code
With astorlm: pattern: 'PLAN_EXECUTE' adds the plan tools and puts the plan in every
request; getPlan() reads it back. The review is plain code after run(), and a second
run() on the same agent continues the same session.
From scratch: A list, two tools that edit it, and a system prompt rebuilt with the list on every
turn. The history lives outside run(), so the fix continues the same conversation.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
const SUBSCRIBERS = [12, 14, 18]
const porches = new Set<number>() // the real world: which porches have a paper
const deliver = tool({
name: 'deliver',
description: 'Ride to a house and throw today’s paper onto its porch. Says where the paper landed.',
schema: z.object({ house: z.number() }),
execute: async ({ house }) => {
const landed = throwPaper(house) // your code: 'porch' or 'bushes'
if (landed === 'porch') porches.add(house)
return landed === 'porch' ? `Paper on the porch at #${house}.` : `Paper landed in the bushes at #${house}.`
},
})
// PLAN: the agent gets add_plan_item and update_plan_item,
// and the current plan is added to the system prompt on every turn.
const agent = await createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
pattern: 'PLAN_EXECUTE',
systemPrompt: 'You deliver newspapers. Plan every stop before you start, and tick each task as you go.',
tools: [deliver],
maxTurns: 20,
})
// REFLECT: check the work itself before accepting the answer. Deterministic when you can;
// a second model with a rubric when you can't.
const review = (): string[] => SUBSCRIBERS.filter((house) => !porches.has(house)).map((house) => `#${house} has no paper on the porch`)
let answer = await agent.run(`Deliver today’s paper to every subscriber on Tango Street: ${SUBSCRIBERS.join(', ')}.`)
for (let round = 1; round <= 2; round++) {
const problems = review()
if (problems.length === 0) break
// Same agent, same session: it keeps its history and its plan, and fixes what's missing.
answer = await agent.run(`Review found: ${problems.join('; ')}. Fix it.`)
}
console.log(agent.getPlan()) // [{ id: '1', description: 'Deliver to #12', status: 'completed' }, …]
console.log(answer.content)
// Plan and reflect, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
const SUBSCRIBERS = [12, 14, 18]
const porches = new Set<number>() // the real world: which porches have a paper
// 1. The plan: a list the model writes and ticks with two tools.
type Task = { id: number; description: string; status: 'pending' | 'completed' }
const plan: Task[] = []
type ToolFn = (args: Record<string, string | number>) => string
const tools: Record<string, ToolFn> = {
add_plan_item: ({ description }) => {
plan.push({ id: plan.length + 1, description: String(description), status: 'pending' })
return `Task added with ID: ${plan.length}`
},
update_plan_item: ({ id, status }) => {
const task = plan.find((t) => t.id === Number(id))
if (!task) return `No task ${id}`
task.status = status === 'completed' ? 'completed' : 'pending'
return `Task ${id} status updated to ${task.status}.`
},
deliver: ({ house }) => {
const landed = throwPaper(Number(house)) // your code: 'porch' or 'bushes'
if (landed === 'porch') porches.add(Number(house))
return landed === 'porch' ? `Paper on the porch at #${house}.` : `Paper landed in the bushes at #${house}.`
},
}
const toolSchemas = [/* one JSON Schema per tool: add_plan_item(description), update_plan_item(id, status), deliver(house) */]
// 2. The plan goes into the system prompt on EVERY turn, so the model never loses track.
const system = (): string =>
'You deliver newspapers. Plan every stop with add_plan_item before you start, and tick each task as you go.\n' +
(plan.length ? plan.map((t) => `- [${t.status}] ${t.description} (id ${t.id})`).join('\n') : '(no plan yet)')
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// The history lives outside run(), so a second run continues the same session.
const messages: Message[] = []
async function run(prompt: string, maxTurns = 20): Promise<string> {
messages.push({ role: 'user', content: prompt })
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: system() }, ...messages], tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const fn = tools[call.function.name]
const output = fn ? fn(JSON.parse(call.function.arguments)) : `Unknown tool: ${call.function.name}`
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// 3. Reflect: check the work itself before accepting the answer, and send what's missing back.
const review = (): string[] => SUBSCRIBERS.filter((house) => !porches.has(house)).map((house) => `#${house} has no paper on the porch`)
let answer = await run(`Deliver today’s paper to every subscriber on Tango Street: ${SUBSCRIBERS.join(', ')}.`)
for (let round = 1; round <= 2; round++) {
const problems = review()
if (problems.length === 0) break
answer = await run(`Review found: ${problems.join('; ')}. Fix it.`)
}
console.log(answer)
# Plan and reflect, from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
SUBSCRIBERS = [12, 14, 18]
PORCHES = set() # the real world: which porches have a paper
# 1. The plan: a list the model writes and ticks with two tools.
PLAN = []
def add_plan_item(description):
PLAN.append({"id": len(PLAN) + 1, "description": description, "status": "pending"})
return f"Task added with ID: {len(PLAN)}"
def update_plan_item(id, status):
task = next((t for t in PLAN if t["id"] == int(id)), None)
if task is None:
return f"No task {id}"
task["status"] = "completed" if status == "completed" else "pending"
return f"Task {id} status updated to {task['status']}."
def deliver(house):
landed = throw_paper(house) # your code: "porch" or "bushes"
if landed == "porch":
PORCHES.add(house)
return f"Paper on the porch at #{house}."
return f"Paper landed in the bushes at #{house}."
TOOLS = {"add_plan_item": add_plan_item, "update_plan_item": update_plan_item, "deliver": deliver}
TOOL_SCHEMAS = [...] # one JSON Schema per tool: add_plan_item(description), update_plan_item(id, status), deliver(house)
# 2. The plan goes into the system prompt on EVERY turn, so the model never loses track.
def system():
lines = [f"- [{t['status']}] {t['description']} (id {t['id']})" for t in PLAN] or ["(no plan yet)"]
return "You deliver newspapers. Plan every stop with add_plan_item before you start, and tick each task as you go.\n" + "\n".join(lines)
# The history lives outside run(), so a second run continues the same session.
MESSAGES = []
def run(prompt, max_turns=20):
MESSAGES.append({"role": "user", "content": prompt})
for _ in range(max_turns):
request = [{"role": "system", "content": system()}, *MESSAGES]
choice = post("/chat/completions", {"model": LLM["model"], "messages": request, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
MESSAGES.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
MESSAGES.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# 3. Reflect: check the work itself before accepting the answer, and send what's missing back.
def review():
return [f"#{house} has no paper on the porch" for house in SUBSCRIBERS if house not in PORCHES]
answer = run(f"Deliver today's paper to every subscriber on Tango Street: {', '.join(map(str, SUBSCRIBERS))}.")
for _ in range(2):
problems = review()
if not problems:
break
answer = run(f"Review found: {'; '.join(problems)}. Fix it.")
print(answer)
What to watch
- Keep tasks small and checkable. “Deliver to #14” can be verified; “handle the street” can’t. A task you can check is a task the review can catch.
- Let the plan change. Plans meet reality: a street is closed, a customer cancels. The agent should be able to add, drop or reorder tasks, not follow a stale list.
- Review the world, not the plan. The sheet said three ticks. Checking the sheet would have passed. The review has to look at the result itself.
- Cap the rounds. A review that can never pass, or an agent that can’t fix what it finds, loops forever. Two or three rounds, then hand it to a person (level 13).
- Don’t plan a one-liner. Planning costs turns and tokens. For a single lookup, skip it. It pays off when the job has several steps that are easy to lose.