第 2 关
智能体循环
单靠模型本身什么也做不了:它只会生成文本。智能体循环才是把它变成智能体的东西。循环把历史记录交给模型,执行模型申请的工具,把结果交回去,然后一直重复,直到模型说“完成”。
1/17
手风琴褶数:
- user
- assistant
- tool_result
EventBus
问题
鸟妈妈问模型:“你能把猪的堡垒清掉吗?”模型什么都发射不了。它最多只能回复一个工具请求:{ name: "launch_red", input: { angle: 40 } }。
如果你的代码只调用一次模型,对话就到此为止了。你手里只剩一个没人执行的请求,没有答案。如果你手动执行这个工具,下一轮又会碰到同样的问题,因为模型可能还需要另一个工具,然后再一个。
解决方案
一个只有一条退出规则的循环:
- 把完整的历史记录加上可用工具列表发给模型。
- 如果回复以
stopReason: "end_turn"结束,就返回它。这是唯一的正常出口。 - 如果以
"tool_use"结束,就执行每个被申请的工具,把结果以tool_result块的形式追加到历史记录里,然后回到第 1 步。
模型决定做什么;循环负责去做。这种分工是其他所有模式的基础:其余的一切(steering、上下文压缩、子智能体)都挂在这个循环的某个点上。
角色
把循环讲成一段小小的冒险。认识了这些角色之后,就没什么需要再破解的了。
- Astor 循环
- 一个小小的探戈舞者,也是唯一会走动的角色。他把问题带给神谕者,每一发都跑到弹弓前,再把答案带回给鸟妈妈。
- 神谕者 Provider
- 模型。它从不碰弹弓:只听手风琴,然后递回一张纸条。需要工具时是橙色,完成时是金色。
- 班多钮手风琴 messages[]
- 历史记录,每条消息是一褶彩色的风箱。风箱每一圈都在变长,而神谕者每次都要把每一褶都听一遍。那些往神谕者飘去的音符就是这个意思。
- 长椅 ToolRegistry
- 每个工具一只鸟:
launch_red、launch_bomb,还有一只今天没人需要。Astor 发射纸条上写的那只,再举起结果:成功就是绿色。 - 鸟妈妈 agent.run()
- 你的代码。它提出问题,然后等待。
- 轨迹、分数和小鸟 历史记录、tokens、maxTurns
- 每一发都在天空留下轨迹,就像历史记录保留每个结果一样。分数就是 token,而且因为整个历史记录都要重新发送,每一圈都涨得更多。每一圈都要从顶栏那排小鸟里扣掉一只,那就是轮数预算。数字仅作示意。
EventBus 面板展示真实循环在动画每一步发出的事件。
代码
使用 astorlm:同一个循环就在 src/agent/loop.ts 里,带有流式输出、重试、钩子、并行执行工具和取消功能。从外面看,它是这样的。
从零手写:大约 40 行代码,可对接任何兼容 OpenAI 的端点,不用 SDK:TypeScript 里只用 fetch,Python 里只用标准库。上面的三个步骤都在注释里标出来了。把开头的 LLM 配置块换成你自己的端点、模型和密钥。
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
// Your game's functions, wrapped as tools: one per bird.
const angle = z.number().min(10).max(80).describe('Launch angle in degrees')
const launchRed = tool({
name: 'launch_red',
description: 'Fling the red bird. Good against wood. Returns what fell and how many pigs are left.',
schema: z.object({ angle }),
execute: async ({ angle }) => level.fling('red', angle), // your code
})
const launchBomb = tool({
name: 'launch_bomb',
description: 'Fling the bomb bird. It explodes on impact: the one to use against stone.',
schema: z.object({ angle }),
execute: async ({ angle }) => level.fling('bomb', angle), // your code
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
tools: [launchRed, launchBomb],
maxTurns: 10, // the birds in line: a cap on the laps
})
agent.on('tool-start', (tool) => console.log('→', tool.name, tool.input))
const answer = await agent.run('The pigs took our eggs! Can you clear their fort?')
// Agent loop from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
type ToolFn = (args: Record<string, string>) => Promise<string>
const tools: Record<string, ToolFn> = { launch_red: launchRed, launch_bomb: launchBomb }
const toolSchemas = [/* one JSON Schema per tool */]
export async function runAgent(prompt: string, maxTurns = 10): Promise<string> {
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
// 1. Send the whole history plus the tool list.
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
// 2. No tool calls: the model is done. The only normal exit.
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
// 3. Run each requested tool and feed the result back as a message.
for (const call of reply.tool_calls ?? []) {
const run = tools[call.function.name]
let output = `Unknown tool: ${call.function.name}`
if (run) {
try {
output = await run(JSON.parse(call.function.arguments))
} catch (err) {
output = `Error: ${err instanceof Error ? err.message : err}`
}
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
# Agent loop from scratch. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
TOOLS = {"launch_red": launch_red, "launch_bomb": launch_bomb}
TOOL_SCHEMAS = [...] # one JSON Schema per tool
def chat(messages):
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)["choices"][0]
def run_agent(prompt, max_turns=10):
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
# 1. Send the whole history plus the tool list.
choice = chat(messages)
reply = choice["message"]
messages.append(reply)
# 2. No tool calls: the model is done. The only normal exit.
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
# 3. Run each requested tool and feed the result back as a message.
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
run = TOOLS.get(name)
try:
output = run(**json.loads(call["function"]["arguments"])) if run else f"Unknown tool: {name}"
except Exception as err:
output = f"Error: {err}"
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
注意,工具出错并不会让循环停下:错误会以文本的形式回到模型那里,让它在下一轮自我纠正。
何时使用,以及要注意什么
只要模型需要在回答之前先行动(读取、搜索、运行某些东西),就用它。如果你只需要转换文本,一次调用就够了,而且更便宜。
- 成本每一圈都在增长。留意手风琴和 token 计数器:神谕者每一轮都要把每一褶都听一遍,所以每一圈都会把之前的一切重新发送一次。轮数一多,上下文就会被填满。
-
一定要设置
maxTurns。一个犯糊涂的模型可能会永远不停地申请工具。没有上限,循环也会永远跑下去。