第 10 关
每圈重新开始
- user
- assistant
- tool_result
EventBus
问题
有些工作一次运行装不下:迁移 300 个文件、翻译整个目录、修好一个仓库里每个失败的测试。每一步都会往历史记录里加一个工具调用和一个结果,而循环每一轮都要把这些全部重新发送。
这就是 Muddle,没完没了的会话。到了下午,它背着从早上开始的每一步:请求庞大无比,旧结果把新结果埋了起来,模型开始重做已经做过的工作,或者跳过只是计划过的工作。什么都没崩溃。质量只是在一点点流失。
压缩(第 7 关)能让 Muddle 慢下来,却拦不住它:工作够长的话,最后会去总结它自己的总结。
解决方案
不要让一个智能体从头活到尾。分圈来跑。每一圈,你的代码都启动一个历史记录为空、目标相同的新智能体。它完成一片工作,写下目前的进展,然后结束。接着你的代码检查工作本身,如果还没完成,就开始下一圈。
-
一次长会话
整项工作都用同一个智能体、同一段历史记录。
每一轮都要把从头开始的一切重新发送。请求越来越重,模型越来越难从中找到要紧的东西,超过窗口就崩了。
-
边做边压缩
还是同一个会话,但在历史记录接近上限时缩小旧消息(第 7 关)。
能争取时间,但解决不了问题。每次压缩都会丢失细节,工作够长的话,最后会压缩到它自己的总结。
-
每圈重新开始
把工作分成若干圈。每一圈都是一个历史记录为空的新智能体。它需要知道的东西,都从文件里读。
每一圈都从小而干净的状态开始。代价是:每一圈都要花一两轮来摸清情况,而且文件里必须写清所有要紧的事。
诀窍在于:任何重要的东西都不放在历史记录里。工作成果在磁盘上(那座桥),说明做到哪一步的简短笔记也在磁盘上(PROGRESS.md)。新的智能体不需要记得上一圈。它只需要读。
这个模式常被叫作 Ralph loop,得名于一个一行 shell 命令:它一遍又一遍地把同一个提示词喂给编程智能体。编程智能体会在长时间的重构中使用它,用 git 工作树和一个 TODO 文件作为状态。
角色
还是那群熟悉的角色,这次换到了峡谷里。
- 舱口 你的代码
- 每一圈启动一个新的智能体(
createIterationAgent),并拿回它的答案。它是包在循环外面的循环。 - 一圈里的 Astor 一次智能体运行
- 第 2 关的智能体循环,带着自己的手风琴。它从空开始,这一圈结束时就飘走。
- 桥 工作成果
- 工具在磁盘上改动的东西。没有哪一圈会把它扔掉。
- 告示牌 PROGRESS.md
- 每一圈留给下一圈的简短字条:哪些做完了,接下来做什么。
- DONE? isDone
- 你的检查,在两圈之间进行。它测量的是桥,而不是模型怎么说。
- LAP 3/5 maxIterations
- 保险丝。如果工作始终通不过检查,循环照样会停下。
看看顶部的两根条。这一圈是每次请求真正的分量,每一圈都重新开始。1 次会话是如果由一个智能体跑完全部三圈,同样的请求会有多重:它永远不会下降。
代码
使用 astorlm:runGoalLoop 接收一个返回新智能体的工厂函数、你的 isDone 检查,以及一个 maxIterations 保险丝。工具写入文件,所以每一圈都能在上一圈留下的地方找到工作成果。
从零手写:第 2 关的循环,放在一个 for 里调用。历史记录是每次调用的局部变量,所以每一圈都天然从空开始。
import { OpenAIProvider, createLocalAgent, runGoalLoop, tool } from 'astorlm'
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// The state lives on disk, not in any history: the bridge, and a progress note.
const GAP = 36
const bridgeLength = (): number => (existsSync('bridge.json') ? JSON.parse(readFileSync('bridge.json', 'utf8')).length : 0)
const readProgress = tool({
name: 'read_progress',
description: 'Read PROGRESS.md: what earlier laps built, and where to start.',
schema: z.object({}),
execute: async () => (existsSync('PROGRESS.md') ? readFileSync('PROGRESS.md', 'utf8') : 'Nothing built yet.'),
})
const layBricks = tool({
name: 'lay_bricks',
description: 'Lay up to 12 bricks of the bridge, starting at brick number "from".',
schema: z.object({ from: z.number().int().min(1), count: z.number().int().min(1).max(12) }),
execute: async ({ from, count }) => {
const to = Math.min(from + count - 1, GAP)
writeFileSync('bridge.json', JSON.stringify({ length: Math.max(bridgeLength(), to) }))
return `Laid bricks ${from}-${to}. The bridge is ${bridgeLength()} bricks long.`
},
})
const writeProgress = tool({
name: 'write_progress',
description: 'Overwrite PROGRESS.md with where the bridge stands now, for whoever comes next.',
schema: z.object({ text: z.string() }),
execute: async ({ text }) => {
writeFileSync('PROGRESS.md', `# Progress\n${text}\n`)
return 'Saved PROGRESS.md.'
},
})
const result = await runGoalLoop({
goal: 'Build the bridge to the exit: 36 bricks. Read PROGRESS.md first, lay at most 12 bricks, then update PROGRESS.md.',
// A NEW agent every lap: empty history, fresh context window. Same tools, same folder.
createIterationAgent: () =>
createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
tools: [readProgress, layBricks, writeProgress],
maxTurns: 8,
}),
// Your code decides when the job is done, by checking the work itself. Not the model's word.
isDone: () => bridgeLength() >= GAP,
onIteration: ({ iteration, lastText }) => console.log(`lap ${iteration}: ${lastText}`),
maxIterations: 5, // the fuse: a goal that never checks out can't run forever
})
console.log(result) // { iterations: 3, done: true, stopReason: 'done', lastText: '…' }
// Fresh laps, from scratch. Plain fetch and node:fs, no SDK.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. The state lives on disk: the bridge, and a progress note for the next lap.
const GAP = 36
const bridgeLength = (): number => (existsSync('bridge.json') ? JSON.parse(readFileSync('bridge.json', 'utf8')).length : 0)
type ToolFn = (args: Record<string, string | number>) => string
const tools: Record<string, ToolFn> = {
read_progress: () => (existsSync('PROGRESS.md') ? readFileSync('PROGRESS.md', 'utf8') : 'Nothing built yet.'),
lay_bricks: ({ from, count }) => {
const to = Math.min(Number(from) + Math.min(Number(count), 12) - 1, GAP)
writeFileSync('bridge.json', JSON.stringify({ length: Math.max(bridgeLength(), to) }))
return `Laid bricks ${from}-${to}. The bridge is ${bridgeLength()} bricks long.`
},
write_progress: ({ text }) => {
writeFileSync('PROGRESS.md', `# Progress\n${text}\n`)
return 'Saved PROGRESS.md.'
},
}
const toolSchemas = [/* one JSON Schema per tool: read_progress(), lay_bricks(from, count), write_progress(text) */]
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 2. The loop from level 2, unchanged. `messages` is born and dies inside each call.
async function runAgent(prompt: string, maxTurns = 8): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
try {
if (run) output = run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// 3. The goal loop: a fresh run per lap, then YOUR check of the work on disk.
const GOAL = 'Build the bridge to the exit: 36 bricks. Read PROGRESS.md first, lay at most 12 bricks, then update PROGRESS.md.'
const MAX_LAPS = 5 // the fuse
for (let lap = 1; lap <= MAX_LAPS; lap++) {
console.log(`lap ${lap}:`, await runAgent(GOAL))
if (bridgeLength() >= GAP) {
console.log(`Done after ${lap} laps.`)
break
}
if (lap === MAX_LAPS) throw new Error(`Bridge unfinished after ${MAX_LAPS} laps: ${bridgeLength()}/${GAP}`)
}
# Fresh laps, from scratch. Standard library only, no SDK.
import json
import urllib.request
from pathlib import Path
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# 1. The state lives on disk: the bridge, and a progress note for the next lap.
GAP = 36
BRIDGE = Path("bridge.json")
PROGRESS = Path("PROGRESS.md")
def bridge_length():
return json.loads(BRIDGE.read_text())["length"] if BRIDGE.exists() else 0
def read_progress():
return PROGRESS.read_text() if PROGRESS.exists() else "Nothing built yet."
def lay_bricks(start, count):
end = min(start + min(count, 12) - 1, GAP)
BRIDGE.write_text(json.dumps({"length": max(bridge_length(), end)}))
return f"Laid bricks {start}-{end}. The bridge is {bridge_length()} bricks long."
def write_progress(text):
PROGRESS.write_text(f"# Progress\n{text}\n")
return "Saved PROGRESS.md."
TOOLS = {
"read_progress": read_progress,
"lay_bricks": lambda **args: lay_bricks(args["from"], args["count"]), # "from" is a Python keyword
"write_progress": write_progress,
}
TOOL_SCHEMAS = [...] # one JSON Schema per tool: read_progress(), lay_bricks(from, count), write_progress(text)
# 2. The loop from level 2, unchanged. `messages` is born and dies inside each call.
def run_agent(prompt, max_turns=8):
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
try:
output = TOOLS[call["function"]["name"]](**json.loads(call["function"]["arguments"]))
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# 3. The goal loop: a fresh run per lap, then YOUR check of the work on disk.
GOAL = "Build the bridge to the exit: 36 bricks. Read PROGRESS.md first, lay at most 12 bricks, then update PROGRESS.md."
MAX_LAPS = 5 # the fuse
for lap in range(1, MAX_LAPS + 1):
print(f"lap {lap}:", run_agent(GOAL))
if bridge_length() >= GAP:
print(f"Done after {lap} laps.")
break
else:
raise RuntimeError(f"Bridge unfinished after {MAX_LAPS} laps: {bridge_length()}/{GAP}")
注意事项
-
检查工作,而不是答案。模型说一句“完成了!”证明不了任何事。
isDone应该看结果本身:跑测试、数行数、量桥长。让它便宜且确定,因为每一圈之后它都要跑一次。 -
一定要装保险丝。一个永远通不过的检查,或者一个不停推翻自己工作的智能体,会一直转下去,直到你的账单让它停下。设置
maxIterations,并看看它为什么会用尽。 - 进度文件是唯一的交接。它漏写的东西,下一圈就不会知道。明确告诉智能体要在里面写什么:哪些做完了、接下来做什么、试过哪些失败了。
- 让每一步都可以安全地重复。一圈可能在中途死掉,工作做完了,字条却还没写。下一圈会把那一片再做一遍,所以做两次绝不能弄坏任何东西。
- 把每一片切小。一圈应该能在一次短运行里完成。如果单独一片就已经需要压缩,说明切得太大了。