Level 9
Memory
- user
- assistant
- tool_result
EventBus
The problem
Everything an agent knows about a conversation lives in its history, the messages the loop resends every turn. When the session ends, the history goes with it. Tomorrow the same customer comes back, says “the same as last time”, and the agent has no idea what that was.
That’s the Eraser Wraith, session amnesia. It doesn’t break anything. It just makes the agent ask the same questions every day, forget the preferences it was told, and treat a regular like a stranger.
Keeping the whole history forever doesn’t fix it either. It grows without end, and Gulp from level 7 is waiting for it.
The solution
Save what matters outside the history, where the end of a session can’t reach it: a file, a database. Then give the agent a way to get it back. There are three common ways, and real agents often mix them:
-
Resume the session
Save the whole history under a session id, and load it back the next time.
Nothing gets lost, but everything comes back: the next request starts heavy, and old chatter crowds the context. Good for picking up an unfinished task, not for remembering a customer for months.
-
Notes in the prompt
The agent saves short notes to a file, and the whole file goes into the system prompt at the start of every session.
Simple and predictable: the model always sees every note. It stops scaling once the notes outgrow a page. Claude Code’s CLAUDE.md and memory files work this way.
-
Search by meaning
Each note is saved with its embedding. A recall tool embeds the question and returns only the closest notes.
Scales to thousands of notes, and finds them even when the words don’t match. Costs an embedding call per note and per search, and the model has to think of calling recall.
The animation shows the third one. An embedding is a list of numbers a model computes for a
text, so that texts that mean similar things get similar numbers. “What the customer ordered last time” and “Buys 1
kg Colombia” share no words, but their embeddings point the same way, and that’s how recall finds the
right page.
The cast
Same cast as always, on a farm this time.
- A day a session
- One conversation, from the first request to the answer. Night ends it.
- The bandoneón the history
- Every message of today’s session. It starts empty every morning.
- The diary long-term memory
-
Notes saved outside any session, one line per note.
rememberwrites a page,recallsearches them. - The Eraser Wraith session end
- Comes every night and empties the bandoneón. It can’t touch the diary.
- The ranking similarity
- How close each note’s meaning is to the query, from 0 to 1. The model only gets the top ones.
- The roaster place_order
- An ordinary tool.
In the EventBus panel, remember and recall are ordinary tool calls. The loop knows
nothing about memory: it’s your tools, and a file they write to.
The code
With astorlm: createSemanticIndex with createOpenAIEmbedder ranks the
notes by meaning. The index lives in memory, so the remember tool also writes it to a file, and the
next session loads it back with addVector.
From scratch: An embedding call, a cosine similarity, a JSON file, and two tools. The loop from level 2 doesn’t change.
import { OpenAIProvider, createLocalAgent, createOpenAIEmbedder, createSemanticIndex, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// The diary: one file of notes per customer, each saved with its embedding.
// The semantic index lives in memory, so load the saved vectors into it at startup.
const customerId = 'c-2291'
const file = `./memory/${customerId}.json`
const diary = createSemanticIndex({
embedder: createOpenAIEmbedder({ ...LLM, model: 'your-embedding-model' }), // e.g. 'nomic-embed-text'
})
if (existsSync(file)) JSON.parse(readFileSync(file, 'utf8')).forEach(diary.addVector)
const remember = tool({
name: 'remember',
description: 'Save a short note about this customer for future conversations.',
schema: z.object({ note: z.string().describe('One fact, in your own words.') }),
execute: async ({ note }) => {
await diary.add(`note-${diary.size + 1}`, note) // embeds it, then stores it
writeFileSync(file, JSON.stringify(diary.list())) // outlives the session
return `Saved. ${diary.size} notes about this customer.`
},
})
const recall = tool({
name: 'recall',
description: 'Search the saved notes about this customer by meaning. Returns the closest ones.',
schema: z.object({ query: z.string() }),
execute: async ({ query }) => {
const hits = await diary.query(query, { topK: 2, threshold: 0.3 })
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
},
})
const agent = await createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
systemPrompt:
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.',
tools: [placeOrder, remember, recall], // placeOrder: your code
maxTurns: 10,
})
// A new session every day: the history starts empty, the diary doesn't.
const last = await agent.run('Hi again! Send me the same as last time.')
console.log(last.content)
// Long-term memory, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
embeddingModel: 'your-embedding-model', // e.g. 'nomic-embed-text', 'text-embedding-3-small'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. An embedding: a list of numbers that captures what a text means.
// Texts that mean similar things get vectors that point the same way.
async function embed(text: string): Promise<number[]> {
const res = await fetch(`${LLM.baseURL}/embeddings`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.embeddingModel, input: text }),
})
return (await res.json()).data[0].embedding
}
// How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
function cosine(a: number[], b: number[]): number {
let dot = 0, na = 0, nb = 0
for (let i = 0; i < a.length; i++) {
dot += a[i]! * b[i]!
na += a[i]! ** 2
nb += b[i]! ** 2
}
return dot / (Math.sqrt(na) * Math.sqrt(nb))
}
// 2. The diary: a file per customer, outside any session. Each note keeps its vector.
type Note = { text: string; vector: number[] }
const FILE = './memory/c-2291.json'
const diary: Note[] = existsSync(FILE) ? JSON.parse(readFileSync(FILE, 'utf8')) : []
// 3. Two tools: one writes a note, the other searches by meaning.
async function remember({ note }: { note: string }): Promise<string> {
diary.push({ text: note, vector: await embed(note) })
writeFileSync(FILE, JSON.stringify(diary))
return `Saved. ${diary.length} notes about this customer.`
}
async function recall({ query }: { query: string }): Promise<string> {
const q = await embed(query)
const hits = diary
.map((note) => ({ text: note.text, score: cosine(q, note.vector) }))
.sort((a, b) => b.score - a.score)
.slice(0, 2)
.filter((hit) => hit.score > 0.3)
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
}
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = {
remember: (args) => remember({ note: args.note ?? '' }),
recall: (args) => recall({ query: args.query ?? '' }),
place_order: placeOrder, // your code
}
const toolSchemas = [/* one JSON Schema per tool: remember(note), recall(query), place_order(…) */]
const SYSTEM =
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.'
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 4. The loop from level 2. Every call is a new session: the history starts empty.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: SYSTEM },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// Two days, two sessions. Nothing but the diary carries over.
await runAgent('Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.')
console.log(await runAgent('Hi again! Send me the same as last time.'))
# Long-term memory, from scratch. Standard library only, no SDK.
import json
import math
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"embedding_model": "your-embedding-model", # e.g. "nomic-embed-text", "text-embedding-3-small"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# 1. An embedding: a list of numbers that captures what a text means.
# Texts that mean similar things get vectors that point the same way.
def embed(text):
return post("/embeddings", {"model": LLM["embedding_model"], "input": text})["data"][0]["embedding"]
# How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
# 2. The diary: a file per customer, outside any session. Each note keeps its vector.
FILE = Path("memory/c-2291.json")
DIARY = json.loads(FILE.read_text()) if FILE.exists() else []
# 3. Two tools: one writes a note, the other searches by meaning.
def remember(note):
DIARY.append({"text": note, "vector": embed(note)})
FILE.write_text(json.dumps(DIARY))
return f"Saved. {len(DIARY)} notes about this customer."
def recall(query):
q = embed(query)
hits = sorted(({"text": n["text"], "score": cosine(q, n["vector"])} for n in DIARY), key=lambda h: -h["score"])
lines = [f"{h['score']:.2f} {h['text']}" for h in hits[:2] if h["score"] > 0.3]
return "\n".join(lines) or "Nothing saved about that."
TOOLS = {"remember": remember, "recall": recall, "place_order": place_order} # place_order: your code
TOOL_SCHEMAS = [...] # one JSON Schema per tool: remember(note), recall(query), place_order(...)
SYSTEM = (
"You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping "
"(what they buy, how they like it), save it with remember. If they refer to the past, use recall first."
)
# 4. The loop from level 2. Every call is a new session: the history starts empty.
def run_agent(prompt, max_turns=10):
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
try:
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# Two days, two sessions. Nothing but the diary carries over.
run_agent("Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.")
print(run_agent("Hi again! Send me the same as last time."))
What to watch
- Decide what’s worth keeping. Save facts that will matter next time: preferences, decisions, addresses. Not the whole chat. Say it in the system prompt, or the model will either save nothing or save everything.
- Keep memories apart. One diary per customer, and check whose it is on every call. A memory search across customers is a data leak waiting to happen.
- Memories go stale. The customer moves from Rosario to Córdoba. Save a date with each note, let newer notes win, and give people a way to see and delete what’s stored about them.
- A match isn’t proof. A search always returns its closest notes, even when none of them fits. Set a minimum score, and when the best match is weak, have the model ask instead of guessing.
- What it reads, it may obey. A saved note goes back into the prompt later. Never let one customer’s text become instructions for another session (level 14).