> Agent Harness Patterns 第 0 关，一条讲解 AI 智能体工作原理的模式路线。网页版：https://harnesspatterns.dev/zh/patterns/your-toolkit · 全部模式（英文）：https://harnesspatterns.dev/llms.txt

# 你的工具箱

在任何智能体之前，先有四个零件：一个只会读写文本的模型，一个告诉它该扮演谁的 system prompt，一个由你的代码每次重发的消息列表，以及它可以申请调用的工具。这张地图后面的一切都由它们搭建而成。

## 四个零件

- 模型 `model`
   角色
   读文本，写文本。它从训练中学到了很多，但对你的应用、你的用户和今天发生的事一无所知。
- system prompt `system`
   装备
   放在每次请求最前面的指令：它是谁、它的规则、它的语气，以及它自己无从得知的事实。
- 消息 `messages[]`
   背包
   到目前为止的对话。你的代码保存这个列表，并在每次调用时完整发送。
- 工具 `tools`
   技能
   描述模型可以申请调用的函数的卡片。它只能申请：真正执行的是你的代码。

## 模型什么都不记得

这一点最让人意外。模型在两次调用之间没有任何记忆。每次请求都从零开始，它只知道这次请求里装着的东西：system prompt、消息和工具列表。

在动画里，第二个问题单独到达，神谕者问“哪个城市？”，尽管刚刚才有人告诉过它。只有当你的代码把之前的消息重新发过去，它才“记得”。聊天应用看起来有记忆，是因为它们每次都把整段对话重新发送一遍。

由此可以得出两点。历史记录归你管：由你来保存、裁剪和存储。而你保留的每条消息，每次调用都会再发送一次，所以对话越长，每次的成本就越高。

## 文本，还是一个请求

当它的列表里有工具时，回复可以是两种之一：给用户的文本，或者用某些输入调用某个工具的请求。模型从不执行任何东西。它写下 `get_forecast(city, date)` 就停下了；执行它是你的代码的工作。

注意那个日期：模型把“明天”换算成了 `2026-09-26`，因为 system prompt 告诉了它今天是几号。它自己无从得知的事实，就该放在那里。

## 代码

**使用 astorlm：**一个 astorlm 智能体拥有同样的四个零件。它会在多次 `run()` 调用之间替你保存历史记录；当模型申请调用工具时，它会执行工具并把结果送回去。这个循环就是第 2 关。

**从零手写：**四个零件加一次请求，还没有循环。把 `LLM` 配置块换成你自己的端点、模型和密钥。

**使用 astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

// A tool: the card the model reads (name, description, schema) plus your code behind it.
const getForecast = tool({
  name: 'get_forecast',
  description: 'Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.',
  schema: z.object({ city: z.string(), date: z.string() }),
  execute: async ({ city, date }) => forecastLine(city, date), // your code; the model never sees it
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  systemPrompt: 'You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line.',
  contextFiles: [], // by default astorlm also appends AGENTS.md and CLAUDE.md from the working folder
  tools: [getForecast],
})

// The agent keeps the history for you, so the second run knows about the first.
await agent.run('I’m in Buenos Aires.')
await agent.run('Will it rain tomorrow?') // asks for get_forecast(Buenos Aires, 2026-09-26), runs it, answers
console.log(agent.getMessages().length) // every message so far, resent on every call
```

**TypeScript**

```ts
// The four pieces, with no agent yet: one request in, one reply out. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }

// 1. The system prompt: who it is, its rules, and facts it can't know on its own.
const system = 'You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line.'

// 2. The history. The model remembers nothing between calls: this array IS its memory.
const messages: Message[] = []

// 3. A tool, as the model sees it: a name, a description and its inputs. Never the code.
const tools = [
  {
    type: 'function',
    function: {
      name: 'get_forecast',
      description: 'Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.',
      parameters: {
        type: 'object',
        properties: { city: { type: 'string' }, date: { type: 'string' } },
        required: ['city', 'date'],
      },
    },
  },
]

// 4. The model: every call sends ALL of the above, and gets back one message.
async function ask(text: string): Promise<Message> {
  messages.push({ role: 'user', content: text })
  const res = await fetch(`${LLM.baseURL}/chat/completions`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
    body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: system }, ...messages], tools }),
  })
  const reply: Message = (await res.json()).choices[0].message
  messages.push(reply) // keep it, or the next call won't know it happened
  return reply
}

await ask('I’m in Buenos Aires.') // "Got it! How can I help?"
const reply = await ask('Will it rain tomorrow?')
// The reply is either text (reply.content) or a tool request (reply.tool_calls):
// get_forecast({ city: "Buenos Aires", date: "2026-09-26" })
// Running it and sending the result back, in a loop, is level 2.
```

**Python**

```python
# The four pieces, with no agent yet: one request in, one reply out. Standard library only, no SDK.
import json
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

# 1. The system prompt: who it is, its rules, and facts it can't know on its own.
SYSTEM = "You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line."

# 2. The history. The model remembers nothing between calls: this list IS its memory.
messages = []

# 3. A tool, as the model sees it: a name, a description and its inputs. Never the code.
TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_forecast",
            "description": "Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}, "date": {"type": "string"}},
                "required": ["city", "date"],
            },
        },
    }
]

# 4. The model: every call sends ALL of the above, and gets back one message.
def ask(text):
    messages.append({"role": "user", "content": text})
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({
            "model": LLM["model"],
            "messages": [{"role": "system", "content": SYSTEM}, *messages],
            "tools": TOOLS,
        }).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        reply = json.load(response)["choices"][0]["message"]
    messages.append(reply)  # keep it, or the next call won't know it happened
    return reply

ask("I'm in Buenos Aires.")  # "Got it! How can I help?"
reply = ask("Will it rain tomorrow?")
# The reply is either text (reply["content"]) or a tool request (reply["tool_calls"]):
# get_forecast(city="Buenos Aires", date="2026-09-26")
# Running it and sending the result back, in a loop, is level 2.
```

## 注意事项

- **让 system prompt 简短而具体。**每次调用它都会跟着发送。写清角色、规则、语气和模型需要的事实，不要写成一本手册。
- **决定历史记录保留什么。**永远把一切都重新发送，会变慢、变贵，最后还会装不下。裁剪和总结历史记录本身就是一个模式。
- **没有工具的模型照样会回答。**问它实时数据，它会流畅地瞎猜。如果答案取决于它看不到的东西，就给它一个工具。

## 相关模式

- [1 · 什么是智能体？](https://harnesspatterns.dev/zh/patterns/what-is-an-agent.md)
- [2 · 智能体循环](https://harnesspatterns.dev/zh/patterns/agent-loop.md)
- [3 · 设计一个工具](https://harnesspatterns.dev/zh/patterns/designing-a-tool.md)
- [7 · 背包装满了](https://harnesspatterns.dev/zh/patterns/compaction.md)
