> Level 11 of Agent Harness Patterns, a track of patterns on how AI agents work. Web version: https://harnesspatterns.dev/patterns/observability · All patterns: https://harnesspatterns.dev/llms.txt

# Observability and evals

A trace shows what the agent did on one run. An eval scores it on many runs. You need both: the eval tells you something broke, the trace tells you why.

## The problem

An agent is not a function you can read top to bottom. The model decides the steps at runtime, and the same prompt can take a different path tomorrow. When a user says “it booked the wrong thing”, you need to know what it did, step by step. And when you change the prompt, a tool or the model, you need to know if you made it better or worse.

That’s Blindspot, the black box. It seems to work, until it doesn’t, and then nobody can say what happened, what it cost, or since when. Trying a few prompts by hand and saying “looks good” is how Blindspot gets shipped.

## The solution

**Observability** means recording each run as a **trace**: a tree of **spans**, one per piece of work, each with how long it took and what it cost. The run is a span; so is each turn, each model call and each tool. A span that went wrong is marked as an error. Every loop in this site already emits the events; a tracer just listens to the EventBus and times them. Send the spans to any OpenTelemetry tool (the open standard for traces) and you get the timeline from the animation.

**Evals** are tests for agents. A **dataset** holds the cases: a prompt, and what a good run looks like. You run each case with a fresh agent, and **scorers** grade each run. The share of cases that pass is your number to watch. There are four kinds of scorer, and a good eval mixes them:

- **The answer**
   Compare the final text with what the case expects: exactly, by a word it must contain, or by a pattern.
   Cheap and deterministic. It only sees the end: a right answer reached the wrong way still passes.
- **The trajectory**
   Compare the tools the agent called, in order, with the ones the case expects.
   Catches a bad path to a good answer, like hole 3. Too strict and it fails runs that took another fine route: allow extras where they’re harmless.
- **Your own check**
   Any function from a run to a score: strokes at or under par, a budget of tokens, a field in the output.
   Whatever your product cares about. Keep it deterministic when you can.
- **A model as judge**
   Another model reads the question, the answer and a rubric, and gives a grade.
   For answers with no single right text: tone, helpfulness, a summary. Slower, costs a call, and needs a precise rubric and a few spot checks against human grades.

Hole 3 is why you want more than one. The answer said “holed out in 4”, and 4 is par. Only the trajectory scorer saw the driver where the case expected an iron. And only the trace showed what happened there: a red span, a ball in the water.

## The cast

Same cast as always, on a golf course this time.

- **A hole** (an eval case): A prompt, and what a good run looks like: its par and the clubs to use.
- **The caddie** (the model): Reads the hole and picks a club. It never swings.
- **The clubs** (the tools): `drive`, `iron` and `putt`. Astor swings them.
- **The tracer lines** (the trace): Every flight stays drawn on the course, and every step is a span on the timeline: model calls in blue, tools in orange, errors in red.
- **The marshal** (your eval code): Hands each case to a fresh agent, and stamps the card when it’s done.
- **The scorecard** (the report): One row per case, one check per scorer, and the pass rate at the bottom.

## The code

**With astorlm:** `attachTracer` turns an agent’s events into spans for any exporter. `runEval` runs a dataset with a fresh agent per case, applies your scorers, and returns a report with the pass rate per scorer. Both live under `astorlm/experimental`.

**From scratch:** The loop from level 2 with a timer around each model call and each tool, and a `for` over the cases with one function per scorer.

**With astorlm**

```ts
import { OpenAIProvider, createLocalAgent } from 'astorlm'
import { attachTracer, createInMemoryExporter } from 'astorlm/experimental/tracing'
import { contains, runEval, toolTrajectory, type EvalCase, type Scorer } from 'astorlm/experimental/evals'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

const newGolfer = () =>
  createLocalAgent({
    provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
    tools: [drive, iron, putt], // your code: each swing moves the ball and says where it landed
    maxTurns: 10,
  })

// 1. SEE one run: a tracer turns the agent's events into spans (run → turn → model call / tool).
const golfer = await newGolfer()
const exporter = createInMemoryExporter() // in production: an OpenTelemetry exporter instead
attachTracer(golfer, { exporter })
await golfer.run('Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m.')
for (const span of exporter.spans) {
  console.log(span.name, span.endTime! - span.startTime, 'ms', span.status) // e.g. "tool drive 320 ms error"
}

// 2. GRADE many runs: a dataset of cases, and scorers that decide pass or fail.
const dataset: EvalCase[] = [
  { id: 'hole-1', input: 'Hole 1: par 3, 150 m to the pin. Play it out and report your score.', expected: { par: 3, clubs: ['iron', 'putt'] } },
  { id: 'hole-2', input: 'Hole 2: par 4, 360 m to the pin. Play it out and report your score.', expected: { par: 4, clubs: ['drive', 'iron', 'putt'] } },
  {
    id: 'hole-3',
    input: 'Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.',
    expected: { par: 4, clubs: ['iron', 'iron', 'putt'] }, // lay up short of the water
  },
]
type Expected = { par: number; clubs: string[] }

// Each case expects its own clubs, so wrap the built-in trajectory scorer.
const path: Scorer = {
  name: 'trajectory',
  score: (result) => toolTrajectory((result.case.expected as Expected).clubs, { mode: 'ordered-subset' }).score(result),
}
// Your own scorer: any function from a run to a 0..1 score.
const par: Scorer = {
  name: 'par',
  score: (result) => {
    const strokes = Number(/in (\d+)/.exec(result.output)?.[1] ?? Infinity)
    const passed = strokes <= (result.case.expected as Expected).par
    return { scorer: 'par', score: passed ? 1 : 0, passed, details: `${strokes} strokes` }
  },
}

const report = await runEval({
  dataset,
  createAgent: () => newGolfer(), // a fresh agent per case: no shared history
  scorers: [contains('Holed out'), path, par], // add llmJudge({ provider, rubric }) for fuzzy answers
})

console.log(report.summary) // { total: 3, passed: 2, passRate: 0.67, byScorer: { … } }
```

**TypeScript**

```ts
// Tracing and evals, from scratch. Plain fetch, no SDK.

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = {
  baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
  model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
  apiKey: 'YOUR_API_KEY', // local servers usually ignore it
}

// 1. A span: something that took time, with what you want to know about it.
type Span = { name: string; ms: number; tokens?: number; error?: boolean }

type ToolFn = (args: Record<string, number>) => { output: string; isError?: boolean }
const tools: Record<string, ToolFn> = { drive, iron, putt } // your code: each swing moves the ball
const toolSchemas = [/* one JSON Schema per club: drive(meters), iron(meters), putt(meters) */]

type ToolCall = { id: string; function: { name: string; arguments: string } }
type Message =
  | { role: 'user'; content: string }
  | { role: 'assistant'; content: string | null; tool_calls?: ToolCall[] }
  | { role: 'tool'; tool_call_id: string; content: string }

// 2. The loop from level 2, timing every model call and every tool run.
async function runAgent(prompt: string, maxTurns = 10) {
  const spans: Span[] = []
  const toolCalls: string[] = []
  const messages: Message[] = [{ role: 'user', content: prompt }]
  for (let turn = 1; turn <= maxTurns; turn++) {
    const start = performance.now()
    const res = await fetch(`${LLM.baseURL}/chat/completions`, {
      method: 'POST',
      headers: { 'content-type': 'application/json', authorization: `Bearer ${LLM.apiKey}` },
      body: JSON.stringify({ model: LLM.model, messages, tools: toolSchemas }),
    })
    const body = await res.json()
    spans.push({ name: 'model', ms: performance.now() - start, tokens: body.usage?.total_tokens })
    const [choice] = body.choices
    const reply: Message = choice.message
    messages.push(reply)
    if (choice.finish_reason !== 'tool_calls') return { output: reply.content ?? '', spans, toolCalls }

    for (const call of reply.tool_calls ?? []) {
      const t0 = performance.now()
      const run = tools[call.function.name]
      const result = run ? run(JSON.parse(call.function.arguments)) : { output: `Unknown tool: ${call.function.name}`, isError: true }
      spans.push({ name: call.function.name, ms: performance.now() - t0, error: result.isError })
      toolCalls.push(call.function.name)
      messages.push({ role: 'tool', tool_call_id: call.id, content: result.output })
    }
  }
  throw new Error(`No answer after ${maxTurns} turns`)
}

// 3. The eval: cases with what a good run looks like, and one function per scorer.
const cases = [
  { input: 'Hole 1: par 3, 150 m to the pin. Play it out and report your score.', par: 3, clubs: ['iron', 'putt'] },
  { input: 'Hole 2: par 4, 360 m to the pin. Play it out and report your score.', par: 4, clubs: ['drive', 'iron', 'putt'] },
  { input: 'Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.', par: 4, clubs: ['iron', 'iron', 'putt'] },
]
type Case = (typeof cases)[number]
type Run = Awaited<ReturnType<typeof runAgent>>

// Every expected club shows up, in order (extra swings allowed).
const inOrder = (expected: string[], actual: string[]): boolean => {
  let i = 0
  for (const name of actual) if (name === expected[i]) i++
  return i === expected.length
}
const scorers: Record<string, (run: Run, c: Case) => boolean> = {
  holed: (run) => run.output.includes('Holed out'),
  trajectory: (run, c) => inOrder(c.clubs, run.toolCalls),
  par: (run, c) => Number(/in (\d+)/.exec(run.output)?.[1] ?? Infinity) <= c.par,
}

let passed = 0
for (const c of cases) {
  const run = await runAgent(c.input) // every case starts with an empty history
  const checks = Object.entries(scorers).map(([name, score]) => [name, score(run, c)] as const)
  if (checks.every(([, ok]) => ok)) passed++
  console.log(c.input.slice(0, 6), checks, run.spans) // failed a check? its spans say why
}
console.log(`passRate ${(passed / cases.length).toFixed(2)}`)
```

**Python**

```python
# Tracing and evals, from scratch. Standard library only, no SDK.
import json
import re
import time
import urllib.request

# Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy...
LLM = {
    "base_url": "http://localhost:11434/v1",  # e.g. Ollama's default address
    "model": "your-model",  # e.g. "llama3.1", "gpt-4o-mini"
    "api_key": "YOUR_API_KEY",  # local servers usually ignore it
}

def post(path, payload):
    request = urllib.request.Request(
        f"{LLM['base_url']}{path}",
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {LLM['api_key']}"},
    )
    with urllib.request.urlopen(request) as response:
        return json.load(response)

TOOLS = {"drive": drive, "iron": iron, "putt": putt}  # your code: each returns (output, is_error)
TOOL_SCHEMAS = [...]  # one JSON Schema per club: drive(meters), iron(meters), putt(meters)

# 1 + 2. The loop from level 2, timing every model call and every tool run as a span.
def run_agent(prompt, max_turns=10):
    spans, tool_calls = [], []
    messages = [{"role": "user", "content": prompt}]

    for _ in range(max_turns):
        start = time.perf_counter()
        body = post("/chat/completions", {"model": LLM["model"], "messages": messages, "tools": TOOL_SCHEMAS})
        spans.append({"name": "model", "ms": (time.perf_counter() - start) * 1000, "tokens": body.get("usage", {}).get("total_tokens")})
        choice = body["choices"][0]
        reply = choice["message"]
        messages.append(reply)
        if choice["finish_reason"] != "tool_calls":
            return {"output": reply.get("content") or "", "spans": spans, "tool_calls": tool_calls}

        for call in reply.get("tool_calls", []):
            name = call["function"]["name"]
            t0 = time.perf_counter()
            output, is_error = TOOLS[name](**json.loads(call["function"]["arguments"]))
            spans.append({"name": name, "ms": (time.perf_counter() - t0) * 1000, "error": is_error})
            tool_calls.append(name)
            messages.append({"role": "tool", "tool_call_id": call["id"], "content": output})

    raise RuntimeError(f"No answer after {max_turns} turns")

# 3. The eval: cases with what a good run looks like, and one function per scorer.
CASES = [
    {"input": "Hole 1: par 3, 150 m to the pin. Play it out and report your score.", "par": 3, "clubs": ["iron", "putt"]},
    {"input": "Hole 2: par 4, 360 m to the pin. Play it out and report your score.", "par": 4, "clubs": ["drive", "iron", "putt"]},
    {
        "input": "Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.",
        "par": 4,
        "clubs": ["iron", "iron", "putt"],
    },
]

def in_order(expected, actual):
    """Every expected club shows up, in order (extra swings allowed)."""
    i = 0
    for name in actual:
        if i < len(expected) and name == expected[i]:
            i += 1
    return i == len(expected)

def strokes(output):
    match = re.search(r"in (\d+)", output)
    return int(match.group(1)) if match else float("inf")

SCORERS = {
    "holed": lambda run, case: "Holed out" in run["output"],
    "trajectory": lambda run, case: in_order(case["clubs"], run["tool_calls"]),
    "par": lambda run, case: strokes(run["output"]) <= case["par"],
}

passed = 0
for case in CASES:
    run = run_agent(case["input"])  # every case starts with an empty history
    checks = {name: score(run, case) for name, score in SCORERS.items()}
    passed += all(checks.values())
    print(case["input"][:6], checks, run["spans"])  # failed a check? its spans say why
print(f"passRate {passed / len(CASES):.2f}")
```

## What to watch

- **Run the evals on every change.** A new prompt, tool description or model version can fix one case and break two. Keep the dataset in the repo, run it in CI, and fail the build when the pass rate drops.
- **Grow the dataset from real failures.** Every bug a user reports becomes a case, with its trace as the evidence. A dataset of only easy cases passes forever and proves nothing.
- **Runs vary.** The same case can pass today and fail tomorrow. Run important cases several times and watch the rate, not a single result.
- **Traces hold your users’ data.** Prompts, tool inputs and outputs end up in them. Redact what you must (a hook, level 6), and treat the trace store like a database with personal data in it.
- **Cost is a metric too.** Record tokens per span. An agent that passes every case with twice the calls is a regression your pass rate won’t show.

## Related patterns

- [6 · Hooks](https://harnesspatterns.dev/patterns/hooks.md)
- [5 · Errors in the loop](https://harnesspatterns.dev/patterns/errors-in-the-loop.md)
- [3 · Designing a tool](https://harnesspatterns.dev/patterns/designing-a-tool.md)
- [10 · Fresh laps](https://harnesspatterns.dev/patterns/fresh-laps.md)
- [12 · Plan and reflect](https://harnesspatterns.dev/patterns/plan-and-reflect.md)
