第 9 关
记忆
- user
- assistant
- tool_result
EventBus
问题
智能体对一段对话的全部了解都存在它的历史记录里,也就是循环每一轮都重新发送的那些消息。会话一结束,历史记录也跟着没了。明天同一位顾客回来,说“照上次的”,智能体根本不知道上次是什么。
这就是橡皮幽灵,会话失忆症。它不会弄坏任何东西。它只是让智能体每天都问同样的问题,忘掉别人告诉过它的偏好,把一位老主顾当成陌生人。
把整段历史记录永远保留下来也解决不了问题。它会无止境地增长,而第 7 关的 Gulp 正等着它。
解决方案
把要紧的东西保存在历史记录之外,放在会话结束也够不着的地方:一个文件,一个数据库。然后给智能体一个把它取回来的办法。常见的做法有三种,真实的智能体经常混着用:
-
恢复会话
用一个会话 id 保存整段历史记录,下次再加载回来。
什么都不会丢,但什么都会回来:下一次请求一开始就很重,旧的闲聊挤占着上下文。适合接着做没做完的任务,不适合连续几个月记住一位顾客。
-
写进提示词的笔记
智能体把简短的笔记存进一个文件,每次会话开始时,整个文件都放进 system prompt。
简单、可预期:模型总能看到每一条笔记。一旦笔记超过一页,就扩展不下去了。Claude Code 的 CLAUDE.md 和记忆文件就是这样工作的。
-
按含义搜索
每条笔记都和它的 embedding 一起保存。recall 工具为问题计算 embedding,只返回最接近的几条笔记。
能扩展到成千上万条笔记,即使用词对不上也能找到。每条笔记、每次搜索都要花一次 embedding 调用,而且模型得想得起来去调用 recall。
动画展示的是第三种。embedding 是模型为一段文本计算出的一串数字,含义相近的文本会得到相近的数字。“顾客上次订了什么”和“买 1 公斤哥伦比亚”没有一个词相同,但它们的 embedding 指向同一个方向,recall 就是这样找到正确那一页的。
角色
还是那群熟悉的角色,这次换到了农场上。
- 一天 一次会话
- 一段对话,从第一个请求到给出答案。夜晚让它结束。
- 手风琴 历史记录
- 今天这次会话的每一条消息。每天早上它都是空的。
- 日记 长期记忆
- 存在任何会话之外的笔记,一条笔记一行。
remember写下一页,recall在其中搜索。 - 橡皮幽灵 会话结束
- 每晚都来,把手风琴清空。它碰不到日记。
- 排名 相似度
- 每条笔记的含义与查询有多接近,从 0 到 1。模型只拿到排在最前面的几条。
- 烘焙机 place_order
- 一个普通工具。
在 EventBus 面板里,remember 和 recall 都是普通的工具调用。循环对记忆一无所知:那只是你的工具,加上它们写入的一个文件。
代码
使用 astorlm:用 createSemanticIndex 搭配 createOpenAIEmbedder 按含义给笔记排序。索引存在内存里,所以 remember 工具还会把它写进一个文件,下一次会话再用 addVector 加载回来。
从零手写:一次 embedding 调用、一个余弦相似度、一个 JSON 文件和两个工具。第 2 关的循环不用改。
import { OpenAIProvider, createLocalAgent, createOpenAIEmbedder, createSemanticIndex, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// The diary: one file of notes per customer, each saved with its embedding.
// The semantic index lives in memory, so load the saved vectors into it at startup.
const customerId = 'c-2291'
const file = `./memory/${customerId}.json`
const diary = createSemanticIndex({
embedder: createOpenAIEmbedder({ ...LLM, model: 'your-embedding-model' }), // e.g. 'nomic-embed-text'
})
if (existsSync(file)) JSON.parse(readFileSync(file, 'utf8')).forEach(diary.addVector)
const remember = tool({
name: 'remember',
description: 'Save a short note about this customer for future conversations.',
schema: z.object({ note: z.string().describe('One fact, in your own words.') }),
execute: async ({ note }) => {
await diary.add(`note-${diary.size + 1}`, note) // embeds it, then stores it
writeFileSync(file, JSON.stringify(diary.list())) // outlives the session
return `Saved. ${diary.size} notes about this customer.`
},
})
const recall = tool({
name: 'recall',
description: 'Search the saved notes about this customer by meaning. Returns the closest ones.',
schema: z.object({ query: z.string() }),
execute: async ({ query }) => {
const hits = await diary.query(query, { topK: 2, threshold: 0.3 })
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
},
})
const agent = await createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
systemPrompt:
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.',
tools: [placeOrder, remember, recall], // placeOrder: your code
maxTurns: 10,
})
// A new session every day: the history starts empty, the diary doesn't.
const last = await agent.run('Hi again! Send me the same as last time.')
console.log(last.content)
// Long-term memory, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
embeddingModel: 'your-embedding-model', // e.g. 'nomic-embed-text', 'text-embedding-3-small'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. An embedding: a list of numbers that captures what a text means.
// Texts that mean similar things get vectors that point the same way.
async function embed(text: string): Promise<number[]> {
const res = await fetch(`${LLM.baseURL}/embeddings`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.embeddingModel, input: text }),
})
return (await res.json()).data[0].embedding
}
// How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
function cosine(a: number[], b: number[]): number {
let dot = 0, na = 0, nb = 0
for (let i = 0; i < a.length; i++) {
dot += a[i]! * b[i]!
na += a[i]! ** 2
nb += b[i]! ** 2
}
return dot / (Math.sqrt(na) * Math.sqrt(nb))
}
// 2. The diary: a file per customer, outside any session. Each note keeps its vector.
type Note = { text: string; vector: number[] }
const FILE = './memory/c-2291.json'
const diary: Note[] = existsSync(FILE) ? JSON.parse(readFileSync(FILE, 'utf8')) : []
// 3. Two tools: one writes a note, the other searches by meaning.
async function remember({ note }: { note: string }): Promise<string> {
diary.push({ text: note, vector: await embed(note) })
writeFileSync(FILE, JSON.stringify(diary))
return `Saved. ${diary.length} notes about this customer.`
}
async function recall({ query }: { query: string }): Promise<string> {
const q = await embed(query)
const hits = diary
.map((note) => ({ text: note.text, score: cosine(q, note.vector) }))
.sort((a, b) => b.score - a.score)
.slice(0, 2)
.filter((hit) => hit.score > 0.3)
return hits.map((hit) => `${hit.score.toFixed(2)} ${hit.text}`).join('\n') || 'Nothing saved about that.'
}
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = {
remember: (args) => remember({ note: args.note ?? '' }),
recall: (args) => recall({ query: args.query ?? '' }),
place_order: placeOrder, // your code
}
const toolSchemas = [/* one JSON Schema per tool: remember(note), recall(query), place_order(…) */]
const SYSTEM =
'You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping ' +
'(what they buy, how they like it), save it with remember. If they refer to the past, use recall first.'
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 4. The loop from level 2. Every call is a new session: the history starts empty.
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: SYSTEM },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// Two days, two sessions. Nothing but the diary carries over.
await runAgent('Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.')
console.log(await runAgent('Hi again! Send me the same as last time.'))
# Long-term memory, from scratch. Standard library only, no SDK.
import json
import math
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"embedding_model": "your-embedding-model", # e.g. "nomic-embed-text", "text-embedding-3-small"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# 1. An embedding: a list of numbers that captures what a text means.
# Texts that mean similar things get vectors that point the same way.
def embed(text):
return post("/embeddings", {"model": LLM["embedding_model"], "input": text})["data"][0]["embedding"]
# How closely two vectors point the same way: 1 = same meaning, 0 = unrelated.
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
# 2. The diary: a file per customer, outside any session. Each note keeps its vector.
FILE = Path("memory/c-2291.json")
DIARY = json.loads(FILE.read_text()) if FILE.exists() else []
# 3. Two tools: one writes a note, the other searches by meaning.
def remember(note):
DIARY.append({"text": note, "vector": embed(note)})
FILE.write_text(json.dumps(DIARY))
return f"Saved. {len(DIARY)} notes about this customer."
def recall(query):
q = embed(query)
hits = sorted(({"text": n["text"], "score": cosine(q, n["vector"])} for n in DIARY), key=lambda h: -h["score"])
lines = [f"{h['score']:.2f} {h['text']}" for h in hits[:2] if h["score"] > 0.3]
return "\n".join(lines) or "Nothing saved about that."
TOOLS = {"remember": remember, "recall": recall, "place_order": place_order} # place_order: your code
TOOL_SCHEMAS = [...] # one JSON Schema per tool: remember(note), recall(query), place_order(...)
SYSTEM = (
"You are the shop assistant of a coffee roaster. When a customer tells you something worth keeping "
"(what they buy, how they like it), save it with remember. If they refer to the past, use recall first."
)
# 4. The loop from level 2. Every call is a new session: the history starts empty.
def run_agent(prompt, max_turns=10):
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
try:
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# Two days, two sessions. Nothing but the diary carries over.
run_agent("Hi! A kilo of Colombia, ground for a moka pot, shipped to Rosario.")
print(run_agent("Hi again! Send me the same as last time."))
注意事项
- 决定什么值得保留。保存下次会用得上的事实:偏好、决定、地址。不要保存整段聊天。在 system prompt 里写清楚,否则模型要么什么都不存,要么什么都存。
- 让记忆彼此隔离。每位顾客一本日记,每次调用都要核对这是谁的。跨顾客的记忆搜索,就是一次随时会发生的数据泄露。
- 记忆会过时。顾客从罗萨里奥搬到了科尔多瓦。每条笔记都存上日期,让更新的笔记优先,并给人们一个查看和删除自己相关记录的办法。
- 匹配不等于证据。搜索总会返回最接近的几条笔记,哪怕一条都不合适。设一个最低分数;当最佳匹配也很弱时,让模型去问,而不是去猜。
- 它读到的,它就可能照做。保存的笔记之后会回到提示词里。绝不要让一位顾客的文字变成另一个会话的指令(第 14 关)。