Nivel 9
Memoria
- user
- assistant
- tool_result
EventBus
El problema
Todo lo que un agente sabe de una conversación vive en su historial, los mensajes que el bucle reenvía en cada turno. Cuando la sesión termina, el historial se va con ella. Mañana el mismo cliente vuelve, dice “lo mismo que la última vez”, y el agente no tiene idea de qué fue eso.
Ese es el Espectro Borrador, la amnesia de sesión. No rompe nada. Solo hace que el agente haga las mismas preguntas todos los días, olvide las preferencias que le dijeron y trate a un cliente habitual como a un extraño.
Guardar todo el historial para siempre tampoco lo arregla. Crece sin fin, y Gulp, del nivel 7, lo está esperando.
La solución
Guarda lo que importa fuera del historial, donde el fin de una sesión no puede alcanzarlo: un archivo, una base de datos. Después dale al agente una forma de recuperarlo. Hay tres formas comunes, y los agentes reales suelen mezclarlas:
-
Retomar la sesión
Guardar todo el historial bajo un id de sesión, y volver a cargarlo la próxima vez.
No se pierde nada, pero vuelve todo: el siguiente pedido arranca pesado, y la charla vieja satura el contexto. Sirve para retomar una tarea sin terminar, no para recordar a un cliente durante meses.
-
Notas en el prompt
El agente guarda notas cortas en un archivo, y el archivo completo va en el system prompt al principio de cada sesión.
Simple y predecible: el modelo siempre ve todas las notas. Deja de escalar cuando las notas superan una página. El CLAUDE.md y los archivos de memoria de Claude Code funcionan así.
-
Búsqueda por significado
Cada nota se guarda con su embedding. Una herramienta recall calcula el embedding de la pregunta y devuelve solo las notas más cercanas.
Escala a miles de notas, y las encuentra aunque las palabras no coincidan. Cuesta una llamada de embedding por nota y por búsqueda, y al modelo se le tiene que ocurrir llamar a recall.
La animación muestra la tercera. Un embedding es una lista de números que un modelo calcula para
un texto, de modo que los textos que significan cosas parecidas tengan números parecidos. “Qué pidió el cliente la
última vez” y “Compra 1 kg de Colombia” no comparten palabras, pero sus embeddings apuntan en la misma dirección,
y así es como recall encuentra la página correcta.
El elenco
El mismo elenco de siempre, esta vez en una granja.
- Un día una sesión
- Una conversación, del primer pedido a la respuesta. La noche la termina.
- El bandoneón el historial
- Cada mensaje de la sesión de hoy. Empieza vacío cada mañana.
- El diario memoria a largo plazo
-
Notas guardadas fuera de cualquier sesión, una línea por nota.
rememberescribe una página,recallbusca entre ellas. - El Espectro Borrador fin de sesión
- Viene cada noche y vacía el bandoneón. No puede tocar el diario.
- El ranking similitud
- Qué tan cerca está el significado de cada nota de la consulta, de 0 a 1. El modelo solo recibe las mejores.
- El tostador place_order
- Una herramienta común.
En el panel EventBus, remember y recall son llamadas a herramientas comunes. El bucle no
sabe nada de memoria: son tus herramientas, y un archivo en el que escriben.
El código
Con astorlm: createSemanticIndex con createOpenAIEmbedder ordena las
notas por significado. El índice vive en memoria, así que la herramienta remember también lo escribe
en un archivo, y la siguiente sesión lo vuelve a cargar con addVector.
Desde cero: Una llamada de embedding, una similitud coseno, un archivo JSON y dos herramientas. El bucle del nivel 2 no cambia.
import { OpenAIProvider, createLocalAgent, createOpenAIEmbedder, createSemanticIndex, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// The diary: one file of notes per customer, each saved with its embedding.
// The semantic index lives in memory, so load the saved vectors into it at startup.
const customerId = 'c-2291'
const file = `./memory/${customerId}.json`
const diary = createSemanticIndex({
embedder: createOpenAIEmbedder({ ...LLM, model: 'your-embedding-model' }), // e.g. 'nomic-embed-text'
})
if (existsSync(file)) JSON.parse(readFileSync(file, 'utf8')).forEach(diary.addVector)
const remember = tool({
name: 'remember',
description: 'Save a short note about this customer for future conversations.',
schema: z.object({ note: z.string().describe('One fact, in your own words.') }),
execute: async ({ note }) => {
await diary.add(`note-${diary.size + 1}`, note) // embeds it, then stores it
writeFileSync(file, JSON.stringify(diary.list())) // outlives the session
return `Saved. ${diary.size} notes about this customer.`
},
})
const recall = tool({
name: 'recall',
description: 'Search the saved notes about this customer by meaning. Returns the closest ones.',
schema: z.object({ query: z.string() }),
execute: async ({ query }) => {
const hits = await diary.query(query, { topK: 2, threshold: 0.3 })
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
},
})
const agent = await createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
systemPrompt:
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.',
tools: [placeOrder, remember, recall], // placeOrder: your code
maxTurns: 10,
})
// A new session every day: the history starts empty, the diary doesn't.
const last = await agent.run('Hi again! Send me the same as last time.')
console.log(last.content)
// Long-term memory, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
embeddingModel: 'your-embedding-model', // e.g. 'nomic-embed-text', 'text-embedding-3-small'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. An embedding: a list of numbers that captures what a text means.
// Texts that mean similar things get vectors that point the same way.
async function embed(text: string): Promise<number[]> {
const res = await fetch(`${LLM.baseURL}/embeddings`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.embeddingModel, input: text }),
})
return (await res.json()).data[0].embedding
}
// How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
function cosine(a: number[], b: number[]): number {
let dot = 0, na = 0, nb = 0
for (let i = 0; i < a.length; i++) {
dot += a[i]! * b[i]!
na += a[i]! ** 2
nb += b[i]! ** 2
}
return dot / (Math.sqrt(na) * Math.sqrt(nb))
}
// 2. The diary: a file per customer, outside any session. Each note keeps its vector.
type Note = { text: string; vector: number[] }
const FILE = './memory/c-2291.json'
const diary: Note[] = existsSync(FILE) ? JSON.parse(readFileSync(FILE, 'utf8')) : []
// 3. Two tools: one writes a note, the other searches by meaning.
async function remember({ note }: { note: string }): Promise<string> {
diary.push({ text: note, vector: await embed(note) })
writeFileSync(FILE, JSON.stringify(diary))
return `Saved. ${diary.length} notes about this customer.`
}
async function recall({ query }: { query: string }): Promise<string> {
const q = await embed(query)
const hits = diary
.map((note) => ({ text: note.text, score: cosine(q, note.vector) }))
.sort((a, b) => b.score - a.score)
.slice(0, 2)
.filter((hit) => hit.score > 0.3)
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
}
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = {
remember: (args) => remember({ note: args.note ?? '' }),
recall: (args) => recall({ query: args.query ?? '' }),
place_order: placeOrder, // your code
}
const toolSchemas = [/* one JSON Schema per tool: remember(note), recall(query), place_order(…) */]
const SYSTEM =
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.'
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 4. The loop from level 2. Every call is a new session: the history starts empty.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: SYSTEM },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// Two days, two sessions. Nothing but the diary carries over.
await runAgent('Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.')
console.log(await runAgent('Hi again! Send me the same as last time.'))
# Long-term memory, from scratch. Standard library only, no SDK.
import json
import math
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"embedding_model": "your-embedding-model", # e.g. "nomic-embed-text", "text-embedding-3-small"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# 1. An embedding: a list of numbers that captures what a text means.
# Texts that mean similar things get vectors that point the same way.
def embed(text):
return post("/embeddings", {"model": LLM["embedding_model"], "input": text})["data"][0]["embedding"]
# How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
# 2. The diary: a file per customer, outside any session. Each note keeps its vector.
FILE = Path("memory/c-2291.json")
DIARY = json.loads(FILE.read_text()) if FILE.exists() else []
# 3. Two tools: one writes a note, the other searches by meaning.
def remember(note):
DIARY.append({"text": note, "vector": embed(note)})
FILE.write_text(json.dumps(DIARY))
return f"Saved. {len(DIARY)} notes about this customer."
def recall(query):
q = embed(query)
hits = sorted(({"text": n["text"], "score": cosine(q, n["vector"])} for n in DIARY), key=lambda h: -h["score"])
lines = [f"{h['score']:.2f} {h['text']}" for h in hits[:2] if h["score"] > 0.3]
return "\n".join(lines) or "Nothing saved about that."
TOOLS = {"remember": remember, "recall": recall, "place_order": place_order} # place_order: your code
TOOL_SCHEMAS = [...] # one JSON Schema per tool: remember(note), recall(query), place_order(...)
SYSTEM = (
"You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping "
"(what they buy, how they like it), save it with remember. If they refer to the past, use recall first."
)
# 4. The loop from level 2. Every call is a new session: the history starts empty.
def run_agent(prompt, max_turns=10):
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
try:
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# Two days, two sessions. Nothing but the diary carries over.
run_agent("Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.")
print(run_agent("Hi again! Send me the same as last time."))
Qué vigilar
- Decide qué vale la pena guardar. Guarda datos que van a importar la próxima vez: preferencias, decisiones, direcciones. No todo el chat. Dilo en el system prompt, o el modelo no va a guardar nada o lo va a guardar todo.
- Mantén las memorias separadas. Un diario por cliente, y verifica de quién es en cada llamada. Una búsqueda de memoria que cruza clientes es una fuga de datos esperando a ocurrir.
- Las memorias se vuelven obsoletas. El cliente se muda de Rosario a Córdoba. Guarda una fecha con cada nota, deja que ganen las notas más nuevas, y dale a la gente una forma de ver y borrar lo que se guarda sobre ella.
- Una coincidencia no es una prueba. Una búsqueda siempre devuelve sus notas más cercanas, aunque ninguna sirva. Define un puntaje mínimo y, cuando la mejor coincidencia es débil, haz que el modelo pregunte en lugar de adivinar.
- Lo que lee, puede obedecerlo. Una nota guardada vuelve al prompt más adelante. Nunca dejes que el texto de un cliente se convierta en instrucciones para otra sesión (nivel 14).