> Level 0 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/your-toolkit · All patterns: https://harnesspatterns.dev/llms.txt

# Your toolkit

Before any agent, four pieces. A model that only reads and writes text, a system prompt that tells it who to be, a list of messages that your code resends every time, and tools it can ask for. Everything later in this map is built from these.

## The four pieces

- The model `model`
   The character
   Reads text, writes text. Knows a lot from training, but nothing about your app, your user or today.
- The system prompt `system`
   Equip
   Instructions that sit at the top of every request: who it is, its rules, its tone, and facts it can’t know on its own.
- The messages `messages[]`
   Bag
   The conversation so far. Your code keeps this list and sends all of it on every call.
- The tools `tools`
   Skills
   Cards describing functions the model may ask for. It can only ask: your code does the running.

## The model remembers nothing

This is the one that surprises people. A model has no memory between calls. Every request starts from zero, and the only things it knows are the ones inside that request: the system prompt, the messages, and the tool list.

In the animation, the second question arrives alone and the Oracle asks "which city?", even though it was told a moment ago. It only "remembers" once your code sends the earlier messages again. Chat apps feel like they remember because they resend the whole conversation every single time.

Two things follow from that. The history is yours to keep, trim and store. And every message you keep is sent again on every call, so a long conversation costs more each time.

## Text, or a request

With tools on its list, a reply can be one of two things: text for the user, or a request to call a tool with some input. The model never runs anything. It writes `get_forecast(city, date)` and stops; running it is your code's job.

Notice the date: the model turned "tomorrow" into `2026-09-26` because the system prompt told it what today is. Facts it can't know on its own belong there.

## The code

**With astorlm:** An astorlm agent holds the same four pieces. It keeps the history for you across `run()` calls, and when the model asks for a tool it runs it and sends the result back. That loop is level 2.

**From scratch:** The four pieces and a single request, no loop yet. Fill in the `LLM` block with your own endpoint, model and key.

**With astorlm**

```ts
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

// A tool: the card the model reads (name, description, schema) plus your code behind it.
const getForecast = tool({
  name: 'get_forecast',
  description: 'Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.',
  schema: z.object({ city: z.string(), date: z.string() }),
  execute: async ({ city, date }) => forecastLine(city, date), // your code; the model never sees it
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  systemPrompt: 'You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line.',
  contextFiles: [], // by default astorlm also appends AGENTS.md and CLAUDE.md from the working folder
  tools: [getForecast],
})

// The agent keeps the history for you, so the second run knows about the first.
await agent.run('I’m in Buenos Aires.')
await agent.run('Will it rain tomorrow?') // asks for get_forecast(Buenos Aires, 2026-09-26), runs it, answers
console.log(agent.getMessages().length) // every message so far, resent on every call
```

**TypeScript**

```ts
// The four pieces, with no agent yet: one request in, one reply out. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }

// 1. The system prompt: who it is, its rules, and facts it can't know on its own.
const system = 'You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line.'

// 2. The history. The model remembers nothing between calls: this array IS its memory.
const messages: Message[] = []

// 3. A tool, as the model sees it: a name, a description and its inputs. Never the code.
const tools = [
  {
    type: 'function',
    function: {
      name: 'get_forecast',
      description: 'Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.',
      parameters: {
        type: 'object',
        properties: { city: { type: 'string' }, date: { type: 'string' } },
        required: ['city', 'date'],
      },
    },
  },
]

// 4. The model: every call sends ALL of the above, and gets back one message.
async function ask(text: string): Promise<Message> {
  messages.push({ role: 'user', content: text })
  const res = await fetch(`${LLM.baseURL}/chat/completions`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
    body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: system }, ...messages], tools }),
  })
  const reply: Message = (await res.json()).choices[0].message
  messages.push(reply) // keep it, or the next call won't know it happened
  return reply
}

await ask('I’m in Buenos Aires.') // "Got it! How can I help?"
const reply = await ask('Will it rain tomorrow?')
// The reply is either text (reply.content) or a tool request (reply.tool_calls):
// get_forecast({ city: "Buenos Aires", date: "2026-09-26" })
// Running it and sending the result back, in a loop, is level 2.
```

**Python**

```python
# The four pieces, with no agent yet: one request in, one reply out. Standard library only, no SDK.
import json
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

# 1. The system prompt: who it is, its rules, and facts it can't know on its own.
SYSTEM = "You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line."

# 2. The history. The model remembers nothing between calls: this list IS its memory.
messages = []

# 3. A tool, as the model sees it: a name, a description and its inputs. Never the code.
TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_forecast",
            "description": "Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}, "date": {"type": "string"}},
                "required": ["city", "date"],
            },
        },
    }
]

# 4. The model: every call sends ALL of the above, and gets back one message.
def ask(text):
    messages.append({"role": "user", "content": text})
    request = urllib.request.Request(
        f"{LLM['base_url']}/chat/completions",
        data=json.dumps({
            "model": LLM["model"],
            "messages": [{"role": "system", "content": SYSTEM}, *messages],
            "tools": TOOLS,
        }).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        reply = json.load(response)["choices"][0]["message"]
    messages.append(reply)  # keep it, or the next call won't know it happened
    return reply

ask("I'm in Buenos Aires.")  # "Got it! How can I help?"
reply = ask("Will it rain tomorrow?")
# The reply is either text (reply["content"]) or a tool request (reply["tool_calls"]):
# get_forecast(city="Buenos Aires", date="2026-09-26")
# Running it and sending the result back, in a loop, is level 2.
```

## What to watch

- **Keep the system prompt short and specific.** It rides along on every call. Role, rules, tone, and the facts the model needs; not a manual.
- **Decide what the history keeps.** Resending everything forever gets slow and expensive, and eventually doesn't fit. Trimming and summarizing it is its own pattern.
- **A model without tools will still answer.** Ask it about live data and it will guess, fluently. If the answer depends on something it can't see, give it a tool.

## Related patterns

- [1 · What is an agent?](https://harnesspatterns.dev/patterns/what-is-an-agent.md)
- [2 · Agent Loop](https://harnesspatterns.dev/patterns/agent-loop.md)
- [3 · Designing a tool](https://harnesspatterns.dev/patterns/designing-a-tool.md)
- [7 · The backpack fills up](https://harnesspatterns.dev/patterns/compaction.md)
