Nível 9
Memória
- user
- assistant
- tool_result
EventBus
O problema
Tudo o que um agente sabe sobre uma conversa vive no histórico dele, as mensagens que o loop reenvia a cada turno. Quando a sessão termina, o histórico vai junto. Amanhã o mesmo cliente volta, diz “o mesmo da última vez”, e o agente não faz ideia do que foi isso.
Esse é o Espectro Apagador, a amnésia de sessão. Ele não quebra nada. Só faz o agente repetir as mesmas perguntas todo dia, esquecer as preferências que ouviu e tratar um cliente fiel como um estranho.
Guardar o histórico inteiro para sempre também não resolve. Ele cresce sem fim, e o Gulp, do nível 7, está esperando por ele.
A solução
Salve o que importa fora do histórico, onde o fim de uma sessão não alcança: um arquivo, um banco de dados. Depois dê ao agente um jeito de recuperar isso. Há três jeitos comuns, e agentes de verdade muitas vezes os misturam:
-
Retomar a sessão
Salvar o histórico inteiro sob um id de sessão e carregá-lo de volta na próxima vez.
Nada se perde, mas tudo volta: a próxima requisição já começa pesada, e a conversa antiga lota o contexto. Bom para retomar uma tarefa inacabada, não para lembrar de um cliente por meses.
-
Notas no prompt
O agente salva notas curtas num arquivo, e o arquivo inteiro vai no system prompt no início de cada sessão.
Simples e previsível: o modelo sempre vê todas as notas. Para de escalar quando as notas passam de uma página. O CLAUDE.md e os arquivos de memória do Claude Code funcionam assim.
-
Busca por significado
Cada nota é salva com o seu embedding. Uma ferramenta recall calcula o embedding da pergunta e devolve só as notas mais próximas.
Escala para milhares de notas, e as encontra mesmo quando as palavras não batem. Custa uma chamada de embedding por nota e por busca, e o modelo precisa se lembrar de chamar o recall.
A animação mostra o terceiro. Um embedding é uma lista de números que um modelo calcula para um
texto, de forma que textos com significados parecidos recebam números parecidos. “O que o cliente pediu da última
vez” e “Compra 1 kg de Colômbia” não têm nenhuma palavra em comum, mas os embeddings deles apontam na mesma
direção, e é assim que o recall encontra a página certa.
O elenco
O mesmo elenco de sempre, desta vez numa fazenda.
- Um dia uma sessão
- Uma conversa, do primeiro pedido até a resposta. A noite a encerra.
- O bandoneón o histórico
- Cada mensagem da sessão de hoje. Ele começa vazio toda manhã.
- O diário memória de longo prazo
-
Notas salvas fora de qualquer sessão, uma linha por nota.
rememberescreve uma página,recallbusca entre elas. - O Espectro Apagador fim da sessão
- Vem toda noite e esvazia o bandoneón. Não consegue tocar no diário.
- O ranking similaridade
- O quão perto o significado de cada nota está da consulta, de 0 a 1. O modelo só recebe as melhores.
- O torrador place_order
- Uma ferramenta comum.
No painel EventBus, remember e recall são chamadas de ferramenta comuns. O loop não sabe
nada sobre memória: são as suas ferramentas, e um arquivo em que elas escrevem.
O código
Com astorlm: createSemanticIndex com createOpenAIEmbedder ordena as
notas por significado. O índice vive em memória, então a ferramenta remember também o grava num
arquivo, e a sessão seguinte o carrega de volta com addVector.
Do zero: Uma chamada de embedding, uma similaridade de cosseno, um arquivo JSON e duas ferramentas. O loop do nível 2 não muda.
import { OpenAIProvider, createLocalAgent, createOpenAIEmbedder, createSemanticIndex, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// The diary: one file of notes per customer, each saved with its embedding.
// The semantic index lives in memory, so load the saved vectors into it at startup.
const customerId = 'c-2291'
const file = `./memory/${customerId}.json`
const diary = createSemanticIndex({
embedder: createOpenAIEmbedder({ ...LLM, model: 'your-embedding-model' }), // e.g. 'nomic-embed-text'
})
if (existsSync(file)) JSON.parse(readFileSync(file, 'utf8')).forEach(diary.addVector)
const remember = tool({
name: 'remember',
description: 'Save a short note about this customer for future conversations.',
schema: z.object({ note: z.string().describe('One fact, in your own words.') }),
execute: async ({ note }) => {
await diary.add(`note-${diary.size + 1}`, note) // embeds it, then stores it
writeFileSync(file, JSON.stringify(diary.list())) // outlives the session
return `Saved. ${diary.size} notes about this customer.`
},
})
const recall = tool({
name: 'recall',
description: 'Search the saved notes about this customer by meaning. Returns the closest ones.',
schema: z.object({ query: z.string() }),
execute: async ({ query }) => {
const hits = await diary.query(query, { topK: 2, threshold: 0.3 })
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
},
})
const agent = await createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
systemPrompt:
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.',
tools: [placeOrder, remember, recall], // placeOrder: your code
maxTurns: 10,
})
// A new session every day: the history starts empty, the diary doesn't.
const last = await agent.run('Hi again! Send me the same as last time.')
console.log(last.content)
// Long-term memory, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
embeddingModel: 'your-embedding-model', // e.g. 'nomic-embed-text', 'text-embedding-3-small'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. An embedding: a list of numbers that captures what a text means.
// Texts that mean similar things get vectors that point the same way.
async function embed(text: string): Promise<number[]> {
const res = await fetch(`${LLM.baseURL}/embeddings`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.embeddingModel, input: text }),
})
return (await res.json()).data[0].embedding
}
// How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
function cosine(a: number[], b: number[]): number {
let dot = 0, na = 0, nb = 0
for (let i = 0; i < a.length; i++) {
dot += a[i]! * b[i]!
na += a[i]! ** 2
nb += b[i]! ** 2
}
return dot / (Math.sqrt(na) * Math.sqrt(nb))
}
// 2. The diary: a file per customer, outside any session. Each note keeps its vector.
type Note = { text: string; vector: number[] }
const FILE = './memory/c-2291.json'
const diary: Note[] = existsSync(FILE) ? JSON.parse(readFileSync(FILE, 'utf8')) : []
// 3. Two tools: one writes a note, the other searches by meaning.
async function remember({ note }: { note: string }): Promise<string> {
diary.push({ text: note, vector: await embed(note) })
writeFileSync(FILE, JSON.stringify(diary))
return `Saved. ${diary.length} notes about this customer.`
}
async function recall({ query }: { query: string }): Promise<string> {
const q = await embed(query)
const hits = diary
.map((note) => ({ text: note.text, score: cosine(q, note.vector) }))
.sort((a, b) => b.score - a.score)
.slice(0, 2)
.filter((hit) => hit.score > 0.3)
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
}
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = {
remember: (args) => remember({ note: args.note ?? '' }),
recall: (args) => recall({ query: args.query ?? '' }),
place_order: placeOrder, // your code
}
const toolSchemas = [/* one JSON Schema per tool: remember(note), recall(query), place_order(…) */]
const SYSTEM =
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.'
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 4. The loop from level 2. Every call is a new session: the history starts empty.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: SYSTEM },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// Two days, two sessions. Nothing but the diary carries over.
await runAgent('Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.')
console.log(await runAgent('Hi again! Send me the same as last time.'))
# Long-term memory, from scratch. Standard library only, no SDK.
import json
import math
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"embedding_model": "your-embedding-model", # e.g. "nomic-embed-text", "text-embedding-3-small"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# 1. An embedding: a list of numbers that captures what a text means.
# Texts that mean similar things get vectors that point the same way.
def embed(text):
return post("/embeddings", {"model": LLM["embedding_model"], "input": text})["data"][0]["embedding"]
# How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
# 2. The diary: a file per customer, outside any session. Each note keeps its vector.
FILE = Path("memory/c-2291.json")
DIARY = json.loads(FILE.read_text()) if FILE.exists() else []
# 3. Two tools: one writes a note, the other searches by meaning.
def remember(note):
DIARY.append({"text": note, "vector": embed(note)})
FILE.write_text(json.dumps(DIARY))
return f"Saved. {len(DIARY)} notes about this customer."
def recall(query):
q = embed(query)
hits = sorted(({"text": n["text"], "score": cosine(q, n["vector"])} for n in DIARY), key=lambda h: -h["score"])
lines = [f"{h['score']:.2f} {h['text']}" for h in hits[:2] if h["score"] > 0.3]
return "\n".join(lines) or "Nothing saved about that."
TOOLS = {"remember": remember, "recall": recall, "place_order": place_order} # place_order: your code
TOOL_SCHEMAS = [...] # one JSON Schema per tool: remember(note), recall(query), place_order(...)
SYSTEM = (
"You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping "
"(what they buy, how they like it), save it with remember. If they refer to the past, use recall first."
)
# 4. The loop from level 2. Every call is a new session: the history starts empty.
def run_agent(prompt, max_turns=10):
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
try:
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# Two days, two sessions. Nothing but the diary carries over.
run_agent("Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.")
print(run_agent("Hi again! Send me the same as last time."))
O que observar
- Decida o que vale a pena guardar. Salve fatos que vão importar da próxima vez: preferências, decisões, endereços. Não o chat inteiro. Diga isso no system prompt, ou o modelo vai salvar nada ou salvar tudo.
- Mantenha as memórias separadas. Um diário por cliente, e confira de quem é a cada chamada. Uma busca de memória que cruza clientes é um vazamento de dados esperando para acontecer.
- Memórias ficam desatualizadas. O cliente se muda de Rosário para Córdoba. Salve uma data com cada nota, deixe as notas mais novas vencerem, e dê às pessoas um jeito de ver e apagar o que está guardado sobre elas.
- Uma correspondência não é prova. Uma busca sempre devolve as notas mais próximas, mesmo quando nenhuma serve. Defina uma pontuação mínima e, quando a melhor correspondência for fraca, faça o modelo perguntar em vez de chutar.
- O que ele lê, ele pode obedecer. Uma nota salva volta para o prompt depois. Nunca deixe o texto de um cliente virar instrução para outra sessão (nível 14).