Skip to content
astorlm
← Map

Level 12

Plan and reflect

Write the plan before touching anything, and check the work before accepting the answer. The plan keeps the agent on track; the review catches what it says it did but didn’t.
1/52 Bandoneón folds:
  • user
  • assistant
  • tool_result
A paper round on Tango Street. The kiosk is where the model sits, with the route sheet pinned beside it: the plan. The editor is your code.

EventBus

The problem

Give an agent a job with several parts and it starts with whatever is in front of it. Halfway through, the early steps are far back in the history, and it forgets one. At the end it answers “done!” with total confidence, because nothing forces it to look.

That’s Scatterbrain. Two failures in one: no plan, so steps get lost; no check, so a step that went wrong gets reported as done. On Tango Street the tool said plainly that the paper landed in the bushes. The model read it and ticked the box anyway.

The solution

Plan first. Before the first real action, the agent writes the work down as a list of tasks, and ticks each one as it goes. The trick that makes it work: the whole plan is added to every request, so the model always sees what’s done and what’s left, however long the history gets. In astorlm that’s pattern: 'PLAN_EXECUTE': two tools, add_plan_item and update_plan_item, and the plan in the system prompt on every turn.

Reflect before accepting. A ticked box is only what the model says. When run() returns, your code reviews the result before accepting it. If something is missing, the finding goes back as a new message on the same session: the agent keeps its history and its plan, and fixes only what’s wrong. Cap the number of rounds. There are three ways to review, from strongest to weakest:

  • Check the world

    Your code looks at the result itself: the porches, the database rows, the test suite, the file on disk.

    The best reviewer when you can have it: cheap, exact, and it can’t be talked into anything. It needs the work to be checkable by code.

  • A critic model

    A second call reads the task, the answer and a checklist, and lists what’s wrong or missing.

    For work no code can check: a summary, an email, a plan. It costs a call, it can miss things too, and it needs concrete criteria, not “is this good?”.

  • Ask the agent

    The system prompt tells the agent to re-read its work before it answers.

    Free and sometimes enough. But it’s the same model grading itself, with the same blind spots: it ticked #14 once already.

This is the same idea as an eval from level 11, used at runtime: an eval grades runs after the fact to improve the agent; a review grades this run before the user ever sees it.

The cast

Same cast as always, on a paper round this time.

The kiosk the model
The Oracle, behind the counter. It decides every step, and never rides.
The route sheet the plan
The tasks and their boxes. It flashes gold on every turn: it goes into every request.
Astor’s bike the loop
Rides each tool call out and back, with the history in his bandoneón.
A throw deliver
An ordinary tool. It says where the paper landed.
The editor your code
Hands over the job, and reviews the porches before accepting the answer. Its magnifier is review().

The code

With astorlm: pattern: 'PLAN_EXECUTE' adds the plan tools and puts the plan in every request; getPlan() reads it back. The review is plain code after run(), and a second run() on the same agent continues the same session.

From scratch: A list, two tools that edit it, and a system prompt rebuilt with the list on every turn. The history lives outside run(), so the fix continues the same conversation.

import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

const SUBSCRIBERS = [12, 14, 18]
const porches = new Set<number>() // the real world: which porches have a paper

const deliver = tool({
  name: 'deliver',
  description: 'Ride to a house and throw today’s paper onto its porch. Says where the paper landed.',
  schema: z.object({ house: z.number() }),
  execute: async ({ house }) => {
    const landed = throwPaper(house) // your code: 'porch' or 'bushes'
    if (landed === 'porch') porches.add(house)
    return landed === 'porch' ? `Paper on the porch at #${house}.` : `Paper landed in the bushes at #${house}.`
  },
})

// PLAN: the agent gets add_plan_item and update_plan_item,
// and the current plan is added to the system prompt on every turn.
const agent = await createLocalAgent({
  provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
  pattern: 'PLAN_EXECUTE',
  systemPrompt: 'You deliver newspapers. Plan every stop before you start, and tick each task as you go.',
  tools: [deliver],
  maxTurns: 20,
})

// REFLECT: check the work itself before accepting the answer. Deterministic when you can;
// a second model with a rubric when you can't.
const review = (): string[] => SUBSCRIBERS.filter((house) => !porches.has(house)).map((house) => `#${house} has no paper on the porch`)

let answer = await agent.run(`Deliver today’s paper to every subscriber on Tango Street: ${SUBSCRIBERS.join(', ')}.`)
for (let round = 1; round <= 2; round++) {
  const problems = review()
  if (problems.length === 0) break
  // Same agent, same session: it keeps its history and its plan, and fixes what's missing.
  answer = await agent.run(`Review found: ${problems.join('; ')}. Fix it.`)
}

console.log(agent.getPlan()) // [{ id: '1', description: 'Deliver to #12', status: 'completed' }, …]
console.log(answer.content)

What to watch

  • Keep tasks small and checkable. “Deliver to #14” can be verified; “handle the street” can’t. A task you can check is a task the review can catch.
  • Let the plan change. Plans meet reality: a street is closed, a customer cancels. The agent should be able to add, drop or reorder tasks, not follow a stale list.
  • Review the world, not the plan. The sheet said three ticks. Checking the sheet would have passed. The review has to look at the result itself.
  • Cap the rounds. A review that can never pass, or an agent that can’t fix what it finds, loops forever. Two or three rounds, then hand it to a person (level 13).
  • Don’t plan a one-liner. Planning costs turns and tokens. For a single lookup, skip it. It pays off when the job has several steps that are easy to lose.