> Nível 9 de Agent Harness Patterns, uma trilha de padrões sobre como funcionam os agentes de IA. Versão web: https://harnesspatterns.dev/pt/patterns/memory · Todos os padrões (em inglês): https://harnesspatterns.dev/llms.txt

# Memória

O histórico morre com a sessão. A memória de longo prazo é o que você salvar fora dela, mais um jeito de encontrar isso de novo na próxima vez.

## O problema

Tudo o que um agente sabe sobre uma conversa vive no histórico dele, as mensagens que o loop reenvia a cada turno. Quando a sessão termina, o histórico vai junto. Amanhã o mesmo cliente volta, diz “o mesmo da última vez”, e o agente não faz ideia do que foi isso.

Esse é o Espectro Apagador, a amnésia de sessão. Ele não quebra nada. Só faz o agente repetir as mesmas perguntas todo dia, esquecer as preferências que ouviu e tratar um cliente fiel como um estranho.

Guardar o histórico inteiro para sempre também não resolve. Ele cresce sem fim, e o Gulp, do nível 7, está esperando por ele.

## A solução

Salve o que importa **fora do histórico**, onde o fim de uma sessão não alcança: um arquivo, um banco de dados. Depois dê ao agente um jeito de recuperar isso. Há três jeitos comuns, e agentes de verdade muitas vezes os misturam:

- **Retomar a sessão**
   Salvar o histórico inteiro sob um id de sessão e carregá-lo de volta na próxima vez.
   Nada se perde, mas tudo volta: a próxima requisição já começa pesada, e a conversa antiga lota o contexto. Bom para retomar uma tarefa inacabada, não para lembrar de um cliente por meses.
- **Notas no prompt**
   O agente salva notas curtas num arquivo, e o arquivo inteiro vai no system prompt no início de cada sessão.
   Simples e previsível: o modelo sempre vê todas as notas. Para de escalar quando as notas passam de uma página. O CLAUDE.md e os arquivos de memória do Claude Code funcionam assim.
- **Busca por significado**
   Cada nota é salva com o seu embedding. Uma ferramenta recall calcula o embedding da pergunta e devolve só as notas mais próximas.
   Escala para milhares de notas, e as encontra mesmo quando as palavras não batem. Custa uma chamada de embedding por nota e por busca, e o modelo precisa se lembrar de chamar o recall.

A animação mostra o terceiro. Um **embedding** é uma lista de números que um modelo calcula para um texto, de forma que textos com significados parecidos recebam números parecidos. “O que o cliente pediu da última vez” e “Compra 1 kg de Colômbia” não têm nenhuma palavra em comum, mas os embeddings deles apontam na mesma direção, e é assim que o `recall` encontra a página certa.

## O elenco

O mesmo elenco de sempre, desta vez numa fazenda.

- **Um dia** (uma sessão): Uma conversa, do primeiro pedido até a resposta. A noite a encerra.
- **O bandoneón** (o histórico): Cada mensagem da sessão de hoje. Ele começa vazio toda manhã.
- **O diário** (memória de longo prazo): Notas salvas fora de qualquer sessão, uma linha por nota. `remember` escreve uma página, `recall` busca entre elas.
- **O Espectro Apagador** (fim da sessão): Vem toda noite e esvazia o bandoneón. Não consegue tocar no diário.
- **O ranking** (similaridade): O quão perto o significado de cada nota está da consulta, de 0 a 1. O modelo só recebe as melhores.
- **O torrador** (place_order): Uma ferramenta comum.

No painel EventBus, `remember` e `recall` são chamadas de ferramenta comuns. O loop não sabe nada sobre memória: são as suas ferramentas, e um arquivo em que elas escrevem.

## O código

**Com astorlm:** `createSemanticIndex` com `createOpenAIEmbedder` ordena as notas por significado. O índice vive em memória, então a ferramenta `remember` também o grava num arquivo, e a sessão seguinte o carrega de volta com `addVector`.

**Do zero:** Uma chamada de embedding, uma similaridade de cosseno, um arquivo JSON e duas ferramentas. O loop do nível 2 não muda.

**Com astorlm**

```ts
import { OpenAIProvider, createLocalAgent, createOpenAIEmbedder, createSemanticIndex, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

// The diary: one file of notes per customer, each saved with its embedding.
// The semantic index lives in memory, so load the saved vectors into it at startup.
const customerId = 'c-2291'
const file = `./memory/${customerId}.json`
const diary = createSemanticIndex({
  embedder: createOpenAIEmbedder({ ...LLM, model: 'your-embedding-model' }), // e.g. 'nomic-embed-text'
})
if (existsSync(file)) JSON.parse(readFileSync(file, 'utf8')).forEach(diary.addVector)

const remember = tool({
  name: 'remember',
  description: 'Save a short note about this customer for future conversations.',
  schema: z.object({ note: z.string().describe('One fact, in your own words.') }),
  execute: async ({ note }) => {
    await diary.add(`note-${diary.size + 1}`, note) // embeds it, then stores it
    writeFileSync(file, JSON.stringify(diary.list())) // outlives the session
    return `Saved. ${diary.size} notes about this customer.`
  },
})

const recall = tool({
  name: 'recall',
  description: 'Search the saved notes about this customer by meaning. Returns the closest ones.',
  schema: z.object({ query: z.string() }),
  execute: async ({ query }) => {
    const hits = await diary.query(query, { topK: 2, threshold: 0.3 })
    return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
  },
})

const agent = await createLocalAgent({
  provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
  systemPrompt:
    'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
    '(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.',
  tools: [placeOrder, remember, recall], // placeOrder: your code
  maxTurns: 10,
})

// A new session every day: the history starts empty, the diary doesn't.
const last = await agent.run('Hi again! Send me the same as last time.')
console.log(last.content)
```

**TypeScript**

```ts
// Long-term memory, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  embeddingModel: 'your-embedding-model', // e.g. 'nomic-embed-text', 'text-embedding-3-small'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

// 1. An embedding: a list of numbers that captures what a text means.
//    Texts that mean similar things get vectors that point the same way.
async function embed(text: string): Promise<number[]> {
  const res = await fetch(`${LLM.baseURL}/embeddings`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
    body: JSON.stringify({ model: LLM.embeddingModel, input: text }),
  })
  return (await res.json()).data[0].embedding
}

// How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
function cosine(a: number[], b: number[]): number {
  let dot = 0, na = 0, nb = 0
  for (let i = 0; i < a.length; i++) {
    dot += a[i]! * b[i]!
    na += a[i]! ** 2
    nb += b[i]! ** 2
  }
  return dot / (Math.sqrt(na) * Math.sqrt(nb))
}

// 2. The diary: a file per customer, outside any session. Each note keeps its vector.
type Note = { text: string; vector: number[] }
const FILE = './memory/c-2291.json'
const diary: Note[] = existsSync(FILE) ? JSON.parse(readFileSync(FILE, 'utf8')) : []

// 3. Two tools: one writes a note, the other searches by meaning.
async function remember({ note }: { note: string }): Promise<string> {
  diary.push({ text: note, vector: await embed(note) })
  writeFileSync(FILE, JSON.stringify(diary))
  return `Saved. ${diary.length} notes about this customer.`
}

async function recall({ query }: { query: string }): Promise<string> {
  const q = await embed(query)
  const hits = diary
    .map((note) => ({ text: note.text, score: cosine(q, note.vector) }))
    .sort((a, b) => b.score - a.score)
    .slice(0, 2)
    .filter((hit) => hit.score > 0.3)
  return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
}

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = {
  remember: (args) => remember({ note: args.note ?? '' }),
  recall: (args) => recall({ query: args.query ?? '' }),
  place_order: placeOrder, // your code
}
const toolSchemas = [/* one JSON Schema per tool: remember(note), recall(query), place_order(…) */]

const SYSTEM =
  'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
  '(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.'

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// 4. The loop from level 2. Every call is a new session: the history starts empty.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  const messages: Message[] = [
    { role: 'system', content: SYSTEM },
    { role: 'user', content: prompt },
  ]
  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      try {
        if (run) output = await run(JSON.parse(call.function.arguments))
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// Two days, two sessions. Nothing but the diary carries over.
await runAgent('Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.')
console.log(await runAgent('Hi again! Send me the same as last time.'))
```

**Python**

```python
# Long-term memory, from scratch. Standard library only, no SDK.
import json
import math
import urllib.request
from pathlib import Path

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "embedding_model": "your-embedding-model",  # e.g. "nomic-embed-text", "text-embedding-3-small"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

def post(path, payload):
    request = urllib.request.Request(
        f"{LLM['base_url']}{path}",
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

# 1. An embedding: a list of numbers that captures what a text means.
#    Texts that mean similar things get vectors that point the same way.
def embed(text):
    return post("/embeddings", {"model": LLM["embedding_model"], "input": text})["data"][0]["embedding"]

# How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

# 2. The diary: a file per customer, outside any session. Each note keeps its vector.
FILE = Path("memory/c-2291.json")
DIARY = json.loads(FILE.read_text()) if FILE.exists() else []

# 3. Two tools: one writes a note, the other searches by meaning.
def remember(note):
    DIARY.append({"text": note, "vector": embed(note)})
    FILE.write_text(json.dumps(DIARY))
    return f"Saved. {len(DIARY)} notes about this customer."

def recall(query):
    q = embed(query)
    hits = sorted(({"text": n["text"], "score": cosine(q, n["vector"])} for n in DIARY), key=lambda h: -h["score"])
    lines = [f"{h['score']:.2f} {h['text']}" for h in hits[:2] if h["score"] > 0.3]
    return "\n".join(lines) or "Nothing saved about that."

TOOLS = {"remember": remember, "recall": recall, "place_order": place_order}  # place_order: your code
TOOL_SCHEMAS = [...]  # one JSON Schema per tool: remember(note), recall(query), place_order(...)

SYSTEM = (
    "You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping "
    "(what they buy, how they like it), save it with remember. If they refer to the past, use recall first."
)

# 4. The loop from level 2. Every call is a new session: the history starts empty.
def run_agent(prompt, max_turns=10):
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]

    for _ in range(max_turns):
        choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            try:
                output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
            except Exception as err:
                output = f"Error: {err}"
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# Two days, two sessions. Nothing but the diary carries over.
run_agent("Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.")
print(run_agent("Hi again! Send me the same as last time."))
```

## O que observar

- **Decida o que vale a pena guardar.** Salve fatos que vão importar da próxima vez: preferências, decisões, endereços. Não o chat inteiro. Diga isso no system prompt, ou o modelo vai salvar nada ou salvar tudo.
- **Mantenha as memórias separadas.** Um diário por cliente, e confira de quem é a cada chamada. Uma busca de memória que cruza clientes é um vazamento de dados esperando para acontecer.
- **Memórias ficam desatualizadas.** O cliente se muda de Rosário para Córdoba. Salve uma data com cada nota, deixe as notas mais novas vencerem, e dê às pessoas um jeito de ver e apagar o que está guardado sobre elas.
- **Uma correspondência não é prova.** Uma busca sempre devolve as notas mais próximas, mesmo quando nenhuma serve. Defina uma pontuação mínima e, quando a melhor correspondência for fraca, faça o modelo perguntar em vez de chutar.
- **O que ele lê, ele pode obedecer.** Uma nota salva volta para o prompt depois. Nunca deixe o texto de um cliente virar instrução para outra sessão (nível 14).

## Padrões relacionados

- [0 · Sua caixa de ferramentas](https://harnesspatterns.dev/pt/patterns/your-toolkit.md)
- [7 · A mochila enche](https://harnesspatterns.dev/pt/patterns/compaction.md)
- [8 · Skills sob demanda](https://harnesspatterns.dev/pt/patterns/skills.md)
- [14 · Segurança e sandboxing](https://harnesspatterns.dev/pt/patterns/security.md)
- [16 · Agentes proativos](https://harnesspatterns.dev/pt/patterns/proactive-agents.md)
