第 13 关
人在回路
- user
- assistant
- tool_result
- tool_result(错误)
EventBus
问题
一个没有人看着的智能体,速度很快,直到它做出谁都不想要的事:付错了发票、给所有客户群发了邮件、把最后一枚硬币花在了错误的奖品上。而一个事事都停下来问的智能体,并不比你自己动手更快。
这就是 Autopilot Rex:从不询问,从不等待,从不接电话。它死守第一个计划,无视之后发生的变化,在本该提问的时候瞎猜。
解决方案
在少数几个需要人来判断的地方安排一个人,其余地方都让智能体自己跑。这样的地方有三个,区别在于由谁来打破沉默:
-
审批
智能体等待
在不可撤销的工具运行之前,一个钩子把它将要做的事原原本本地展示给人看,然后等一个“同意”或“拒绝”。
付款、删除、发给客户的消息、最后一枚硬币:任何撤不回来的事。
-
中途纠正
人来打断
人随时可以把一条纠正放进队列。到下一次工具调用时,循环取消这次调用,改为把纠正交给模型。
有人在旁边看着,改了主意,或者发现智能体正朝错误的方向走。
-
上报
智能体来问
一个像 ask_human 这样的工具,在人回答之前不会返回。什么时候用它,由模型决定。
信息不足、两个选项一样好、没有把握。system prompt 写明什么时候该问。
审批和中途纠正都放在第 6 关的 beforeToolExecution 钩子里:它在每个工具之前运行,而且可以等待。审批时,它等待人的回答;中途纠正时,它检查队列里有没有纠正。不管哪种情况,人说的话都会作为工具结果回到模型那里,所以运行会带着新信息继续下去,而不是崩溃。
角色
还是那群熟悉的角色,这次换到了游戏厅里。
- 算命师 模型
- 坐在摊位里的神谕者。它阅读历史记录,决定下一次调用。它从不碰那台机器。
- 摇杆和按钮 工具
move_claw不花钱,可以撤回。drop_claw会花掉最后一枚硬币:撤不回来。- Tina 人
- 她的气泡说明是谁先开口的:! 是她打断,? 是有人问她,YES! 是她同意了。
- 屏幕 审批
- DROP? YES NO:钩子等待时,爪子一动不动。她做出决定之前,什么都不会运行。
在 EventBus 面板里,中途纠正显示为一个 user_steering 事件,后面跟着被取消那次调用的结果。审批和回答没有专门的事件:它们只是一个钩子和一个工具,多花了点时间而已。
代码
使用 astorlm:为不可撤销的工具写一个审批钩子,再用 createSteeringController 包起来,它会加上纠正队列:从你的界面调用 steer(text),纠正就会在下一次工具调用时生效。上报就是一个等待你界面响应的普通工具。
从零手写:在第 2 关循环的工具步骤里加三项检查:队列里的纠正、危险工具的审批,以及一个等待人回应的工具。
import { OpenAIProvider, createLocalAgent, createSteeringController, tool } from 'astorlm'
import { z } from 'zod'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
// Your UI: each of these resolves when the person clicks or types.
declare function approve(what: string): Promise<boolean> // shows YES / NO
declare function answer(question: string): Promise<string> // shows a text box
// 1. ESCALATE: a tool the model calls when it isn't sure. It waits for a person.
const askKid = tool({
name: 'ask_kid',
description: 'Ask Tina when you are not sure what she wants. Waits for her answer.',
schema: z.object({ question: z.string() }),
execute: async ({ question }) => answer(question),
})
// 2. APPROVE: irreversible tools wait for a yes before they run.
const NEEDS_APPROVAL = new Set(['drop_claw'])
const steering = createSteeringController({
beforeToolExecution: async ({ toolName, input }) => {
if (!NEEDS_APPROVAL.has(toolName)) return { authorize: true }
const ok = await approve(`${toolName}(${JSON.stringify(input)})`) // show the real call
return ok ? { authorize: true } : { authorize: false, mockResult: 'Tina said no. Ask her what to do.' }
},
})
const agent = await createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
systemPrompt: 'You work a claw machine for Tina. If you are not sure which prize she means, ask her before acting.',
tools: [moveClaw, dropClaw, askKid], // moveClaw, dropClaw: your code
hooks: steering.hooks, // the steering controller wraps the approval hook
maxTurns: 20,
})
// 3. STEER: the person can correct the agent at any moment, e.g. from a button.
// The loop cancels the next tool call and hands the model this feedback instead.
onTinaShouts((text) => steering.steer(text)) // your UI: 'No, wait! The penguin!'
const result = await agent.run('Get me the bear!')
console.log(result.content)
// Human in the loop, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// Your UI: each of these resolves when the person clicks or types.
declare function approve(what: string): Promise<boolean>
declare function answer(question: string): Promise<string>
type ToolFn = (args: Record<string, string | number>) => Promise<string>
const tools: Record<string, ToolFn> = {
move_claw: moveClaw, // your code
drop_claw: dropClaw, // your code
// 1. ESCALATE: the model asks, a person answers.
ask_kid: ({ question }) => answer(String(question)),
}
const toolSchemas = [/* one JSON Schema per tool: move_claw(to), drop_claw(), ask_kid(question) */]
const NEEDS_APPROVAL = new Set(['drop_claw'])
// 3. STEER: a person can queue a correction at any time, e.g. from a button.
let steer: string | null = null
export const queueCorrection = (text: string) => {
steer = text
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'system' | 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
export async function runAgent(prompt: string, maxTurns = 20): Promise<string> {
const messages: Message[] = [
{ role: 'system', content: 'You work a claw machine for Tina. If you are not sure which prize she means, ask her before acting.' },
{ role: 'user', content: prompt },
]
for (let turn = 1; turn <= maxTurns; turn++) {
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const [choice] = (await res.json()).choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return reply.content ?? ''
for (const call of reply.tool_calls ?? []) {
const name = call.function.name
const args = JSON.parse(call.function.arguments)
let output: string
if (steer !== null) {
// The tool boundary: a queued correction cancels this call, and the model reads it instead.
output = `Cancelled. The person says: ${steer}`
steer = null
} else if (NEEDS_APPROVAL.has(name) && !(await approve(`${name}(${call.function.arguments})`))) {
// 2. APPROVE: irreversible tools wait here for a yes.
output = 'Tina said no. Ask her what to do.'
} else {
output = tools[name] ? await tools[name](args) : `Unknown tool: ${name}`
}
messages.push({ role: 'tool', tool_call_id: call.id, content: output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
console.log(await runAgent('Get me the bear!'))
# Human in the loop, from scratch. Standard library only, no SDK.
import json
import queue
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
# In a terminal the person is input(); in an app, whatever your UI sends back.
def approve(what):
return input(f"Approve {what}? [y/N] ").strip().lower() == "y"
def ask_kid(question): # 1. ESCALATE: the model asks, a person answers.
return input(f"The agent asks: {question}\n> ")
TOOLS = {"move_claw": move_claw, "drop_claw": drop_claw, "ask_kid": ask_kid} # move_claw, drop_claw: your code
TOOL_SCHEMAS = [...] # one JSON Schema per tool: move_claw(to), drop_claw(), ask_kid(question)
NEEDS_APPROVAL = {"drop_claw"}
# 3. STEER: another thread (a UI, a chat) can queue a correction at any time.
CORRECTIONS = queue.Queue()
def run_agent(prompt, max_turns=20):
messages = [
{"role": "system", "content": "You work a claw machine for Tina. If you are not sure which prize she means, ask her before acting."},
{"role": "user", "content": prompt},
]
for _ in range(max_turns):
choice = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return reply.get("content") or ""
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])
if not CORRECTIONS.empty():
# The tool boundary: a queued correction cancels this call, and the model reads it instead.
output = f"Cancelled. The person says: {CORRECTIONS.get()}"
elif name in NEEDS_APPROVAL and not approve(f"{name}({args})"):
# 2. APPROVE: irreversible tools wait here for a yes.
output = "Tina said no. Ask her what to do."
else:
output = TOOLS[name](**args)
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
print(run_agent("Get me the bear!"))
注意事项
- 问得越少,分量越重。只审批不可撤销或昂贵的操作。一个人一小时点二十次 YES,就会不再细看,审批也就沦为走过场。
- 展示真实的调用。只写“要放下吗?”是不够的。把工具和它的确切参数、金额、收件人、奖品都展示出来,让人批准的正是将要真正运行的东西。
- 决定没人回答时怎么办。人可能会走开。设一个超时,并选一个安全的默认值:通常是拒绝,并告诉模型原因。
- 纠正会在下一次工具调用时生效。中途纠正无法让运行到一半的工具停下,也无法让写到一半的回复停在半个词上。如果智能体只是在写文本,纠正就得等着。要强行停下,就中止这次运行。
- 留下记录。记录谁批准了什么、什么时候、当时看到了什么。出了问题时,“是智能体干的”从来都不是故事的全部。