Level 11
Observability and evals
- user
- assistant
- tool_result
- tool_result (error)
EventBus
The problem
An agent is not a function you can read top to bottom. The model decides the steps at runtime, and the same prompt can take a different path tomorrow. When a user says “it booked the wrong thing”, you need to know what it did, step by step. And when you change the prompt, a tool or the model, you need to know if you made it better or worse.
That’s Blindspot, the black box. It seems to work, until it doesn’t, and then nobody can say what happened, what it cost, or since when. Trying a few prompts by hand and saying “looks good” is how Blindspot gets shipped.
The solution
Observability means recording each run as a trace: a tree of spans, one per piece of work, each with how long it took and what it cost. The run is a span; so is each turn, each model call and each tool. A span that went wrong is marked as an error. Every loop in this site already emits the events; a tracer just listens to the EventBus and times them. Send the spans to any OpenTelemetry tool (the open standard for traces) and you get the timeline from the animation.
Evals are tests for agents. A dataset holds the cases: a prompt, and what a good run looks like. You run each case with a fresh agent, and scorers grade each run. The share of cases that pass is your number to watch. There are four kinds of scorer, and a good eval mixes them:
-
The answer
Compare the final text with what the case expects: exactly, by a word it must contain, or by a pattern.
Cheap and deterministic. It only sees the end: a right answer reached the wrong way still passes.
-
The trajectory
Compare the tools the agent called, in order, with the ones the case expects.
Catches a bad path to a good answer, like hole 3. Too strict and it fails runs that took another fine route: allow extras where they’re harmless.
-
Your own check
Any function from a run to a score: strokes at or under par, a budget of tokens, a field in the output.
Whatever your product cares about. Keep it deterministic when you can.
-
A model as judge
Another model reads the question, the answer and a rubric, and gives a grade.
For answers with no single right text: tone, helpfulness, a summary. Slower, costs a call, and needs a precise rubric and a few spot checks against human grades.
Hole 3 is why you want more than one. The answer said “holed out in 4”, and 4 is par. Only the trajectory scorer saw the driver where the case expected an iron. And only the trace showed what happened there: a red span, a ball in the water.
The cast
Same cast as always, on a golf course this time.
- A hole an eval case
- A prompt, and what a good run looks like: its par and the clubs to use.
- The caddie the model
- Reads the hole and picks a club. It never swings.
- The clubs the tools
drive,ironandputt. Astor swings them.- The tracer lines the trace
- Every flight stays drawn on the course, and every step is a span on the timeline: model calls in blue, tools in orange, errors in red.
- The marshal your eval code
- Hands each case to a fresh agent, and stamps the card when it’s done.
- The scorecard the report
- One row per case, one check per scorer, and the pass rate at the bottom.
The code
With astorlm: attachTracer turns an agent’s events into spans for any exporter.
runEval runs a dataset with a fresh agent per case, applies your scorers, and returns a report with the
pass rate per scorer. Both live under astorlm/experimental.
From scratch: The loop from level 2 with a timer around each model call and each tool, and a
for over the cases with one function per scorer.
import { OpenAIProvider, createLocalAgent } from 'astorlm'
import { attachTracer, createInMemoryExporter } from 'astorlm/experimental/tracing'
import { contains, runEval, toolTrajectory, type EvalCase, type Scorer } from 'astorlm/experimental/evals'
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key
const newGolfer = () =>
createLocalAgent({
provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
tools: [drive, iron, putt], // your code: each swing moves the ball and says where it landed
maxTurns: 10,
})
// 1. SEE one run: a tracer turns the agent's events into spans (run → turn → model call / tool).
const golfer = await newGolfer()
const exporter = createInMemoryExporter() // in production: an OpenTelemetry exporter instead
attachTracer(golfer, { exporter })
await golfer.run('Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m.')
for (const span of exporter.spans) {
console.log(span.name, span.endTime! - span.startTime, 'ms', span.status) // e.g. "tool drive 320 ms error"
}
// 2. GRADE many runs: a dataset of cases, and scorers that decide pass or fail.
const dataset: EvalCase[] = [
{ id: 'hole-1', input: 'Hole 1: par 3, 150 m to the pin. Play it out and report your score.', expected: { par: 3, clubs: ['iron', 'putt'] } },
{ id: 'hole-2', input: 'Hole 2: par 4, 360 m to the pin. Play it out and report your score.', expected: { par: 4, clubs: ['drive', 'iron', 'putt'] } },
{
id: 'hole-3',
input: 'Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.',
expected: { par: 4, clubs: ['iron', 'iron', 'putt'] }, // lay up short of the water
},
]
type Expected = { par: number; clubs: string[] }
// Each case expects its own clubs, so wrap the built-in trajectory scorer.
const path: Scorer = {
name: 'trajectory',
score: (result) => toolTrajectory((result.case.expected as Expected).clubs, { mode: 'ordered-subset' }).score(result),
}
// Your own scorer: any function from a run to a 0..1 score.
const par: Scorer = {
name: 'par',
score: (result) => {
const strokes = Number(/in (\d+)/.exec(result.output)?.[1] ?? Infinity)
const passed = strokes <= (result.case.expected as Expected).par
return { scorer: 'par', score: passed ? 1 : 0, passed, details: `${strokes} strokes` }
},
}
const report = await runEval({
dataset,
createAgent: () => newGolfer(), // a fresh agent per case: no shared history
scorers: [contains('Holed out'), path, par], // add llmJudge({ provider, rubric }) for fuzzy answers
})
console.log(report.summary) // { total: 3, passed: 2, passRate: 0.67, byScorer: { … } }
// Tracing and evals, from scratch. Plain fetch, no SDK.
// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}
// 1. A span: something that took time, with what you want to know about it.
type Span = { name: string; ms: number; tokens?: number; error?: boolean }
type ToolFn = (args: Record<string, number>) => { output: string; isError?: boolean }
const tools: Record<string, ToolFn> = { drive, iron, putt } // your code: each swing moves the ball
const toolSchemas = [/* one JSON Schema per club: drive(meters), iron(meters), putt(meters) */]
type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
| { role: 'user'; content: string }
| { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
| { role: 'tool'; tool_call_id: string; content: string }
// 2. The loop from level 2, timing every model call and every tool run.
async function runAgent(prompt: string, maxTurns = 10) {
const spans: Span[] = []
const toolCalls: string[] = []
const messages: Message[] = [{ role: 'user', content: prompt }]
for (let turn = 1; turn <= maxTurns; turn++) {
const start = performance.now()
const res = await fetch(`${LLM.baseURL}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
})
const body = await res.json()
spans.push({ name: 'model', ms: performance.now() - start, tokens: body.usage?.total_tokens })
const [choice] = body.choices
const reply: Message = choice.message
messages.push(reply)
if (choice.finish_reason !== 'tool_calls') return { output: reply.content ?? '', spans, toolCalls }
for (const call of reply.tool_calls ?? []) {
const t0 = performance.now()
const run = tools[call.function.name]
const result = run ? run(JSON.parse(call.function.arguments)) : { output: `Unknown tool: ${call.function.name}`, isError: true }
spans.push({ name: call.function.name, ms: performance.now() - t0, error: result.isError })
toolCalls.push(call.function.name)
messages.push({ role: 'tool', tool_call_id: call.id, content: result.output })
}
}
throw new Error(`No answer after ${maxTurns} turns`)
}
// 3. The eval: cases with what a good run looks like, and one function per scorer.
const cases = [
{ input: 'Hole 1: par 3, 150 m to the pin. Play it out and report your score.', par: 3, clubs: ['iron', 'putt'] },
{ input: 'Hole 2: par 4, 360 m to the pin. Play it out and report your score.', par: 4, clubs: ['drive', 'iron', 'putt'] },
{ input: 'Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.', par: 4, clubs: ['iron', 'iron', 'putt'] },
]
type Case = (typeof cases)[number]
type Run = Awaited<ReturnType<typeof runAgent>>
// Every expected club shows up, in order (extra swings allowed).
const inOrder = (expected: string[], actual: string[]): boolean => {
let i = 0
for (const name of actual) if (name === expected[i]) i++
return i === expected.length
}
const scorers: Record<string, (run: Run, c: Case) => boolean> = {
holed: (run) => run.output.includes('Holed out'),
trajectory: (run, c) => inOrder(c.clubs, run.toolCalls),
par: (run, c) => Number(/in (\d+)/.exec(run.output)?.[1] ?? Infinity) <= c.par,
}
let passed = 0
for (const c of cases) {
const run = await runAgent(c.input) // every case starts with an empty history
const checks = Object.entries(scorers).map(([name, score]) => [name, score(run, c)] as const)
if (checks.every(([, ok]) => ok)) passed++
console.log(c.input.slice(0, 6), checks, run.spans) // failed a check? its spans say why
}
console.log(`passRate ${(passed / cases.length).toFixed(2)}`)
# Tracing and evals, from scratch. Standard library only, no SDK.
import json
import re
import time
import urllib.request
# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
"base_url": "http://localhost:11434/v1", # e.g. Ollama's default address
"model": "your-model", # e.g. "llama3.1", "gpt-4o-mini"
"api_key": "YOUR_API_KEY", # local servers usually ignore it
}
def post(path, payload):
request = urllib.request.Request(
f"{LLM['base_url']}{path}",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
)
with urllib.request.urlopen(request) as response:
return json.load(response)
TOOLS = {"drive": drive, "iron": iron, "putt": putt} # your code: each returns (output, is_error)
TOOL_SCHEMAS = [...] # one JSON Schema per club: drive(meters), iron(meters), putt(meters)
# 1 + 2. The loop from level 2, timing every model call and every tool run as a span.
def run_agent(prompt, max_turns=10):
spans, tool_calls = [], []
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
start = time.perf_counter()
body = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})
spans.append({"name": "model", "ms": (time.perf_counter() - start) * 1000, "tokens": body.get("usage", {}).get("total_tokens")})
choice = body["choices"][0]
reply = choice["message"]
messages.append(reply)
if choice["finish_reason"] != "tool_calls":
return {"output": reply.get("content") or "", "spans": spans, "tool_calls": tool_calls}
for call in reply.get("tool_calls", []):
name = call["function"]["name"]
t0 = time.perf_counter()
output, is_error = TOOLS[name](**json.loads(call["function"]["arguments"]))
spans.append({"name": name, "ms": (time.perf_counter() - t0) * 1000, "error": is_error})
tool_calls.append(name)
messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})
raise RuntimeError(f"No answer after {max_turns} turns")
# 3. The eval: cases with what a good run looks like, and one function per scorer.
CASES = [
{"input": "Hole 1: par 3, 150 m to the pin. Play it out and report your score.", "par": 3, "clubs": ["iron", "putt"]},
{"input": "Hole 2: par 4, 360 m to the pin. Play it out and report your score.", "par": 4, "clubs": ["drive", "iron", "putt"]},
{
"input": "Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.",
"par": 4,
"clubs": ["iron", "iron", "putt"],
},
]
def in_order(expected, actual):
"""Every expected club shows up, in order (extra swings allowed)."""
i = 0
for name in actual:
if i < len(expected) and name == expected[i]:
i += 1
return i == len(expected)
def strokes(output):
match = re.search(r"in (\d+)", output)
return int(match.group(1)) if match else float("inf")
SCORERS = {
"holed": lambda run, case: "Holed out" in run["output"],
"trajectory": lambda run, case: in_order(case["clubs"], run["tool_calls"]),
"par": lambda run, case: strokes(run["output"]) <= case["par"],
}
passed = 0
for case in CASES:
run = run_agent(case["input"]) # every case starts with an empty history
checks = {name: score(run, case) for name, score in SCORERS.items()}
passed += all(checks.values())
print(case["input"][:6], checks, run["spans"]) # failed a check? its spans say why
print(f"passRate {passed / len(CASES):.2f}")
What to watch
- Run the evals on every change. A new prompt, tool description or model version can fix one case and break two. Keep the dataset in the repo, run it in CI, and fail the build when the pass rate drops.
- Grow the dataset from real failures. Every bug a user reports becomes a case, with its trace as the evidence. A dataset of only easy cases passes forever and proves nothing.
- Runs vary. The same case can pass today and fail tomorrow. Run important cases several times and watch the rate, not a single result.
- Traces hold your users’ data. Prompts, tool inputs and outputs end up in them. Redact what you must (a hook, level 6), and treat the trace store like a database with personal data in it.
- Cost is a metric too. Record tokens per span. An agent that passes every case with twice the calls is a regression your pass rate won’t show.