Level 0
Your toolkit
EventBus
The four pieces
-
The model
modelThe character
Reads text, writes text. Knows a lot from training, but nothing about your app, your user or today.
-
The system prompt
systemEquip
Instructions that sit at the top of every request: who it is, its rules, its tone, and facts it can’t know on its own.
-
The messages
messages[]Bag
The conversation so far. Your code keeps this list and sends all of it on every call.
-
The tools
toolsSkills
Cards describing functions the model may ask for. It can only ask: your code does the running.
The model remembers nothing
This is the one that surprises people. A model has no memory between calls. Every request starts from zero, and the only things it knows are the ones inside that request: the system prompt, the messages, and the tool list.
In the animation, the second question arrives alone and the Oracle asks "which city?", even though it was told a moment ago. It only "remembers" once your code sends the earlier messages again. Chat apps feel like they remember because they resend the whole conversation every single time.
Two things follow from that. The history is yours to keep, trim and store. And every message you keep is sent again on every call, so a long conversation costs more each time.
Text, or a request
With tools on its list, a reply can be one of two things: text for the user, or a request to call a tool with
some input. The model never runs anything. It writes get_forecast(city, date) and stops; running it
is your code's job.
Notice the date: the model turned "tomorrow" into 2026-09-26 because the system prompt told it
what today is. Facts it can't know on its own belong there.
The code
With astorlm: An astorlm agent holds the same four pieces. It keeps the history for you across run() calls, and
when the model asks for a tool it runs it and sends the result back. That loop is level 2.
From scratch: The four pieces and a single request, no loop yet. Fill in the LLM block with your own endpoint,
model and key.
import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'
// A tool: the card the model reads (name, description, schema) plus your code behind it.
const getForecast = tool({
name: 'get_forecast',
description: 'Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.',
schema: z.object({ city: z.string(), date: z.string() }),
execute: async ({ city, date }) => forecastLine(city, date), // your code; the model never sees it
})
const agent = await createLocalAgent({
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
provider: new OpenAIProvider({
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}),
systemPrompt: 'You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line.',
contextFiles: [], // by default astorlm also appends AGENTS.md and CLAUDE.md from the working folder
tools: [getForecast],
})
// The agent keeps the history for you, so the second run knows about the first.
await agent.run('I’m in Buenos Aires.')
await agent.run('Will it rain tomorrow?') // asks for get_forecast(Buenos Aires, 2026-09-26), runs it, answers
console.log(agent.getMessages().length) // every message so far, resent on every call
// The four pieces, with no agent yet: one request in, one reply out. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
// 1. The system prompt: who it is, its rules, and facts it can't know on its own.
const system = 'You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line.'
// 2. The history. The model remembers nothing between calls: this array IS its memory.
const messages: Message[] = []
// 3. A tool, as the model sees it: a name, a description and its inputs. Never the code.
const tools = [
{
type: 'function',
function: {
name: 'get_forecast',
description: 'Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.',
parameters: {
type: 'object',
properties: { city: { type: 'string' }, date: { type: 'string' } },
required: ['city', 'date'],
},
},
},
]
// 4. The model: every call sends ALL of the above, and gets back one message.
async function ask(text: string): Promise<Message> {
messages.push({ role: 'user', content: text })
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages: [{ role: 'system', content: system }, ...messages], tools }),
})
const reply: Message = (await res.json()).choices[0].message
messages.push(reply) // keep it, or the next call won't know it happened
return reply
}
await ask('I’m in Buenos Aires.') // "Got it! How can I help?"
const reply = await ask('Will it rain tomorrow?')
// The reply is either text (reply.content) or a tool request (reply.tool_calls):
// get_forecast({ city: "Buenos Aires", date: "2026-09-26" })
// Running it and sending the result back, in a loop, is level 2.
# The four pieces, with no agent yet: one request in, one reply out. Standard library only, no SDK.
import json
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
# 1. The system prompt: who it is, its rules, and facts it can't know on its own.
SYSTEM = "You are Nimbus, a weather assistant. Today is 2026-09-25. Answer in one short line."
# 2. The history. The model remembers nothing between calls: this list IS its memory.
messages = []
# 3. A tool, as the model sees it: a name, a description and its inputs. Never the code.
TOOLS = [
{
"type": "function",
"function": {
"name": "get_forecast",
"description": "Daily forecast for one city: rain chance and min/max temp. date is YYYY-MM-DD.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}, "date": {"type": "string"}},
"required": ["city", "date"],
},
},
}
]
# 4. The model: every call sends ALL of the above, and gets back one message.
def ask(text):
messages.append({"role": "user", "content": text})
request = urllib.request.Request(
f"{LLM['base_url']}/chat/completions",
data=json.dumps({
"model": LLM["model"],
"messages": [{"role": "system", "content": SYSTEM}, *messages],
"tools": TOOLS,
}).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
reply = json.load(response)["choices"][0]["message"]
messages.append(reply) # keep it, or the next call won't know it happened
return reply
ask("I'm in Buenos Aires.") # "Got it! How can I help?"
reply = ask("Will it rain tomorrow?")
# The reply is either text (reply["content"]) or a tool request (reply["tool_calls"]):
# get_forecast(city="Buenos Aires", date="2026-09-26")
# Running it and sending the result back, in a loop, is level 2.
What to watch
- Keep the system prompt short and specific. It rides along on every call. Role, rules, tone, and the facts the model needs; not a manual.
- Decide what the history keeps. Resending everything forever gets slow and expensive, and eventually doesn't fit. Trimming and summarizing it is its own pattern.
- A model without tools will still answer. Ask it about live data and it will guess, fluently. If the answer depends on something it can't see, give it a tool.