> Agent Harness Patterns 第 9 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/memory · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 记忆

历史记录随会话一同消亡。长期记忆就是你存在会话之外的东西，再加上一个下次能把它找回来的办法。

## 问题

智能体对一段对话的全部了解都存在它的历史记录里，也就是循环每一轮都重新发送的那些消息。会话一结束，历史记录也跟着没了。明天同一位顾客回来，说“照上次的”，智能体根本不知道上次是什么。

这就是橡皮幽灵，会话失忆症。它不会弄坏任何东西。它只是让智能体每天都问同样的问题，忘掉别人告诉过它的偏好，把一位老主顾当成陌生人。

把整段历史记录永远保留下来也解决不了问题。它会无止境地增长，而第 7 关的 Gulp 正等着它。

## 解决方案

把要紧的东西保存在**历史记录之外**，放在会话结束也够不着的地方：一个文件，一个数据库。然后给智能体一个把它取回来的办法。常见的做法有三种，真实的智能体经常混着用：

- **恢复会话**
   用一个会话 id 保存整段历史记录，下次再加载回来。
   什么都不会丢，但什么都会回来：下一次请求一开始就很重，旧的闲聊挤占着上下文。适合接着做没做完的任务，不适合连续几个月记住一位顾客。
- **写进提示词的笔记**
   智能体把简短的笔记存进一个文件，每次会话开始时，整个文件都放进 system prompt。
   简单、可预期：模型总能看到每一条笔记。一旦笔记超过一页，就扩展不下去了。Claude Code 的 CLAUDE.md 和记忆文件就是这样工作的。
- **按含义搜索**
   每条笔记都和它的 embedding 一起保存。recall 工具为问题计算 embedding，只返回最接近的几条笔记。
   能扩展到成千上万条笔记，即使用词对不上也能找到。每条笔记、每次搜索都要花一次 embedding 调用，而且模型得想得起来去调用 recall。

动画展示的是第三种。**embedding** 是模型为一段文本计算出的一串数字，含义相近的文本会得到相近的数字。“顾客上次订了什么”和“买 1 公斤哥伦比亚”没有一个词相同，但它们的 embedding 指向同一个方向，`recall` 就是这样找到正确那一页的。

## 角色

还是那群熟悉的角色，这次换到了农场上。

- **一天** (一次会话): 一段对话，从第一个请求到给出答案。夜晚让它结束。
- **手风琴** (历史记录): 今天这次会话的每一条消息。每天早上它都是空的。
- **日记** (长期记忆): 存在任何会话之外的笔记，一条笔记一行。`remember` 写下一页，`recall` 在其中搜索。
- **橡皮幽灵** (会话结束): 每晚都来，把手风琴清空。它碰不到日记。
- **排名** (相似度): 每条笔记的含义与查询有多接近，从 0 到 1。模型只拿到排在最前面的几条。
- **烘焙机** (place_order): 一个普通工具。

在 EventBus 面板里，`remember` 和 `recall` 都是普通的工具调用。循环对记忆一无所知：那只是你的工具，加上它们写入的一个文件。

## 代码

**使用 astorlm：**用 `createSemanticIndex` 搭配 `createOpenAIEmbedder` 按含义给笔记排序。索引存在内存里，所以 `remember` 工具还会把它写进一个文件，下一次会话再用 `addVector` 加载回来。

**从零手写：**一次 embedding 调用、一个余弦相似度、一个 JSON 文件和两个工具。第 2 关的循环不用改。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, createOpenAIEmbedder, createSemanticIndex, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

// The diary: one file of notes per customer, each saved with its embedding.
// The semantic index lives in memory, so load the saved vectors into it at startup.
const customerId = 'c-2291'
const file = `./memory/${customerId}.json`
const diary = createSemanticIndex({
  embedder: createOpenAIEmbedder({ ...LLM, model: 'your-embedding-model' }), // e.g. 'nomic-embed-text'
})
if (existsSync(file)) JSON.parse(readFileSync(file, 'utf8')).forEach(diary.addVector)

const remember = tool({
  name: 'remember',
  description: 'Save a short note about this customer for future conversations.',
  schema: z.object({ note: z.string().describe('One fact, in your own words.') }),
  execute: async ({ note }) => {
    await diary.add(`note-${diary.size + 1}`, note) // embeds it, then stores it
    writeFileSync(file, JSON.stringify(diary.list())) // outlives the session
    return `Saved. ${diary.size} notes about this customer.`
  },
})

const recall = tool({
  name: 'recall',
  description: 'Search the saved notes about this customer by meaning. Returns the closest ones.',
  schema: z.object({ query: z.string() }),
  execute: async ({ query }) => {
    const hits = await diary.query(query, { topK: 2, threshold: 0.3 })
    return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
  },
})

const agent = await createLocalAgent({
  provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
  systemPrompt:
    'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
    '(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.',
  tools: [placeOrder, remember, recall], // placeOrder: your code
  maxTurns: 10,
})

// A new session every day: the history starts empty, the diary doesn't.
const last = await agent.run('Hi again! Send me the same as last time.')
console.log(last.content)
```

**TypeScript**

```ts
// Long-term memory, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  embeddingModel: 'your-embedding-model', // e.g. 'nomic-embed-text', 'text-embedding-3-small'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

// 1. An embedding: a list of numbers that captures what a text means.
//    Texts that mean similar things get vectors that point the same way.
async function embed(text: string): Promise<number[]> {
  const res = await fetch(`${LLM.baseURL}/embeddings`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
    body: JSON.stringify({ model: LLM.embeddingModel, input: text }),
  })
  return (await res.json()).data[0].embedding
}

// How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
function cosine(a: number[], b: number[]): number {
  let dot = 0, na = 0, nb = 0
  for (let i = 0; i < a.length; i++) {
    dot += a[i]! * b[i]!
    na += a[i]! ** 2
    nb += b[i]! ** 2
  }
  return dot / (Math.sqrt(na) * Math.sqrt(nb))
}

// 2. The diary: a file per customer, outside any session. Each note keeps its vector.
type Note = { text: string; vector: number[] }
const FILE = './memory/c-2291.json'
const diary: Note[] = existsSync(FILE) ? JSON.parse(readFileSync(FILE, 'utf8')) : []

// 3. Two tools: one writes a note, the other searches by meaning.
async function remember({ note }: { note: string }): Promise<string> {
  diary.push({ text: note, vector: await embed(note) })
  writeFileSync(FILE, JSON.stringify(diary))
  return `Saved. ${diary.length} notes about this customer.`
}

async function recall({ query }: { query: string }): Promise<string> {
  const q = await embed(query)
  const hits = diary
    .map((note) => ({ text: note.text, score: cosine(q, note.vector) }))
    .sort((a, b) => b.score - a.score)
    .slice(0, 2)
    .filter((hit) => hit.score > 0.3)
  return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
}

type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = {
  remember: (args) => remember({ note: args.note ?? '' }),
  recall: (args) => recall({ query: args.query ?? '' }),
  place_order: placeOrder, // your code
}
const toolSchemas = [/* one JSON Schema per tool: remember(note), recall(query), place_order(…) */]

const SYSTEM =
  'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
  '(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.'

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'system' | 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// 4. The loop from level 2. Every call is a new session: the history starts empty.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
  const messages: Message[] = [
    { role: 'system', content: SYSTEM },
    { role: 'user', content: prompt },
  ]
  for (let turn = 1; turn <= maxTurns; turn++) {
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const [choice] = (await res.json()).choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''

    for (const call of reply.tool_calls ?? []) {
      const run = tools[call.function.name]
      let output = `Unknown tool: ${call.function.name}`
      try {
        if (run) output = await run(JSON.parse(call.function.arguments))
      } catch (err) {
        output = `Error: ${err instanceof Error ? err.message : err}`
      }
      messages.push({ role: 'tool', tool_call_id: call.id, content: output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// Two days, two sessions. Nothing but the diary carries over.
await runAgent('Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.')
console.log(await runAgent('Hi again! Send me the same as last time.'))
```

**Python**

```python
# Long-term memory, from scratch. Standard library only, no SDK.
import json
import math
import urllib.request
from pathlib import Path

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "embedding_model": "your-embedding-model",  # e.g. "nomic-embed-text", "text-embedding-3-small"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

def post(path, payload):
    request = urllib.request.Request(
        f"{LLM['base_url']}{path}",
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

# 1. An embedding: a list of numbers that captures what a text means.
#    Texts that mean similar things get vectors that point the same way.
def embed(text):
    return post("/embeddings", {"model": LLM["embedding_model"], "input": text})["data"][0]["embedding"]

# How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

# 2. The diary: a file per customer, outside any session. Each note keeps its vector.
FILE = Path("memory/c-2291.json")
DIARY = json.loads(FILE.read_text()) if FILE.exists() else []

# 3. Two tools: one writes a note, the other searches by meaning.
def remember(note):
    DIARY.append({"text": note, "vector": embed(note)})
    FILE.write_text(json.dumps(DIARY))
    return f"Saved. {len(DIARY)} notes about this customer."

def recall(query):
    q = embed(query)
    hits = sorted(({"text": n["text"], "score": cosine(q, n["vector"])} for n in DIARY), key=lambda h: -h["score"])
    lines = [f"{h['score']:.2f} {h['text']}" for h in hits[:2] if h["score"] > 0.3]
    return "\n".join(lines) or "Nothing saved about that."

TOOLS = {"remember": remember, "recall": recall, "place_order": place_order}  # place_order: your code
TOOL_SCHEMAS = [...]  # one JSON Schema per tool: remember(note), recall(query), place_order(...)

SYSTEM = (
    "You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping "
    "(what they buy, how they like it), save it with remember. If they refer to the past, use recall first."
)

# 4. The loop from level 2. Every call is a new session: the history starts empty.
def run_agent(prompt, max_turns=10):
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]

    for _ in range(max_turns):
        choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return reply.get("content") or ""

        for call in reply.get("tool_calls", []):
            try:
                output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
            except Exception as err:
                output = f"Error: {err}"
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# Two days, two sessions. Nothing but the diary carries over.
run_agent("Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.")
print(run_agent("Hi again! Send me the same as last time."))
```

## 注意事项

- **决定什么值得保留。**保存下次会用得上的事实：偏好、决定、地址。不要保存整段聊天。在 system prompt 里写清楚，否则模型要么什么都不存，要么什么都存。
- **让记忆彼此隔离。**每位顾客一本日记，每次调用都要核对这是谁的。跨顾客的记忆搜索，就是一次随时会发生的数据泄露。
- **记忆会过时。**顾客从罗萨里奥搬到了科尔多瓦。每条笔记都存上日期，让更新的笔记优先，并给人们一个查看和删除自己相关记录的办法。
- **匹配不等于证据。**搜索总会返回最接近的几条笔记，哪怕一条都不合适。设一个最低分数；当最佳匹配也很弱时，让模型去问，而不是去猜。
- **它读到的，它就可能照做。**保存的笔记之后会回到提示词里。绝不要让一位顾客的文字变成另一个会话的指令（第 14 关）。

## 相关模式

- [0 · 你的工具箱](https://harnesspatterns.dev/zh/patterns/your-toolkit.md)
- [7 · 背包装满了](https://harnesspatterns.dev/zh/patterns/compaction.md)
- [8 · 按需加载的技能](https://harnesspatterns.dev/zh/patterns/skills.md)
- [14 · 安全与沙箱](https://harnesspatterns.dev/zh/patterns/security.md)
- [16 · 主动式智能体](https://harnesspatterns.dev/zh/patterns/proactive-agents.md)
