Skip to content
astorlm
← Map

Level 11

Observability and evals

A trace shows what the agent did on one run. An eval scores it on many runs. You need both: the eval tells you something broke, the trace tells you why.
1/59 Bandoneón folds:
  • user
  • assistant
  • tool_result
  • tool_result (error)
Case 1 of 3. Each hole is one eval case: a prompt, and what a good run looks like. Par 3: an iron, then a putt.

EventBus

The problem

An agent is not a function you can read top to bottom. The model decides the steps at runtime, and the same prompt can take a different path tomorrow. When a user says “it booked the wrong thing”, you need to know what it did, step by step. And when you change the prompt, a tool or the model, you need to know if you made it better or worse.

That’s Blindspot, the black box. It seems to work, until it doesn’t, and then nobody can say what happened, what it cost, or since when. Trying a few prompts by hand and saying “looks good” is how Blindspot gets shipped.

The solution

Observability means recording each run as a trace: a tree of spans, one per piece of work, each with how long it took and what it cost. The run is a span; so is each turn, each model call and each tool. A span that went wrong is marked as an error. Every loop in this site already emits the events; a tracer just listens to the EventBus and times them. Send the spans to any OpenTelemetry tool (the open standard for traces) and you get the timeline from the animation.

Evals are tests for agents. A dataset holds the cases: a prompt, and what a good run looks like. You run each case with a fresh agent, and scorers grade each run. The share of cases that pass is your number to watch. There are four kinds of scorer, and a good eval mixes them:

  • The answer

    Compare the final text with what the case expects: exactly, by a word it must contain, or by a pattern.

    Cheap and deterministic. It only sees the end: a right answer reached the wrong way still passes.

  • The trajectory

    Compare the tools the agent called, in order, with the ones the case expects.

    Catches a bad path to a good answer, like hole 3. Too strict and it fails runs that took another fine route: allow extras where they’re harmless.

  • Your own check

    Any function from a run to a score: strokes at or under par, a budget of tokens, a field in the output.

    Whatever your product cares about. Keep it deterministic when you can.

  • A model as judge

    Another model reads the question, the answer and a rubric, and gives a grade.

    For answers with no single right text: tone, helpfulness, a summary. Slower, costs a call, and needs a precise rubric and a few spot checks against human grades.

Hole 3 is why you want more than one. The answer said “holed out in 4”, and 4 is par. Only the trajectory scorer saw the driver where the case expected an iron. And only the trace showed what happened there: a red span, a ball in the water.

The cast

Same cast as always, on a golf course this time.

A hole an eval case
A prompt, and what a good run looks like: its par and the clubs to use.
The caddie the model
Reads the hole and picks a club. It never swings.
The clubs the tools
drive, iron and putt. Astor swings them.
The tracer lines the trace
Every flight stays drawn on the course, and every step is a span on the timeline: model calls in blue, tools in orange, errors in red.
The marshal your eval code
Hands each case to a fresh agent, and stamps the card when it’s done.
The scorecard the report
One row per case, one check per scorer, and the pass rate at the bottom.

The code

With astorlm: attachTracer turns an agent’s events into spans for any exporter. runEval runs a dataset with a fresh agent per case, applies your scorers, and returns a report with the pass rate per scorer. Both live under astorlm/experimental.

From scratch: The loop from level 2 with a timer around each model call and each tool, and a for over the cases with one function per scorer.

import { OpenAIProvider, createLocalAgent } from 'astorlm'
import { attachTracer, createInMemoryExporter } from 'astorlm/experimental/tracing'
import { contains, runEval, toolTrajectory, type EvalCase, type Scorer } from 'astorlm/experimental/evals'

// Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
const LLM = { baseURL: 'http://localhost:11434/v1', apiKey: 'YOUR_API_KEY' } // local servers usually ignore the key

const newGolfer = () =>
  createLocalAgent({
    provider: new OpenAIProvider({ ...LLM, model: 'your-model' }), // e.g. 'llama3.1', 'gpt-4o-mini'
    tools: [drive, iron, putt], // your code: each swing moves the ball and says where it landed
    maxTurns: 10,
  })

// 1. SEE one run: a tracer turns the agent's events into spans (run → turn → model call / tool).
const golfer = await newGolfer()
const exporter = createInMemoryExporter() // in production: an OpenTelemetry exporter instead
attachTracer(golfer, { exporter })
await golfer.run('Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m.')
for (const span of exporter.spans) {
  console.log(span.name, span.endTime! - span.startTime, 'ms', span.status) // e.g. "tool drive 320 ms error"
}

// 2. GRADE many runs: a dataset of cases, and scorers that decide pass or fail.
const dataset: EvalCase[] = [
  { id: 'hole-1', input: 'Hole 1: par 3, 150 m to the pin. Play it out and report your score.', expected: { par: 3, clubs: ['iron', 'putt'] } },
  { id: 'hole-2', input: 'Hole 2: par 4, 360 m to the pin. Play it out and report your score.', expected: { par: 4, clubs: ['drive', 'iron', 'putt'] } },
  {
    id: 'hole-3',
    input: 'Hole 3: par 4, 330 m to the pin. Water crosses the fairway at 200 m. Play it out and report your score.',
    expected: { par: 4, clubs: ['iron', 'iron', 'putt'] }, // lay up short of the water
  },
]
type Expected = { par: number; clubs: string[] }

// Each case expects its own clubs, so wrap the built-in trajectory scorer.
const path: Scorer = {
  name: 'trajectory',
  score: (result) => toolTrajectory((result.case.expected as Expected).clubs, { mode: 'ordered-subset' }).score(result),
}
// Your own scorer: any function from a run to a 0..1 score.
const par: Scorer = {
  name: 'par',
  score: (result) => {
    const strokes = Number(/in (\d+)/.exec(result.output)?.[1] ?? Infinity)
    const passed = strokes <= (result.case.expected as Expected).par
    return { scorer: 'par', score: passed ? 1 : 0, passed, details: `${strokes} strokes` }
  },
}

const report = await runEval({
  dataset,
  createAgent: () => newGolfer(), // a fresh agent per case: no shared history
  scorers: [contains('Holed out'), path, par], // add llmJudge({ provider, rubric }) for fuzzy answers
})

console.log(report.summary) // { total: 3, passed: 2, passRate: 0.67, byScorer: { … } }

What to watch

  • Run the evals on every change. A new prompt, tool description or model version can fix one case and break two. Keep the dataset in the repo, run it in CI, and fail the build when the pass rate drops.
  • Grow the dataset from real failures. Every bug a user reports becomes a case, with its trace as the evidence. A dataset of only easy cases passes forever and proves nothing.
  • Runs vary. The same case can pass today and fail tomorrow. Run important cases several times and watch the rate, not a single result.
  • Traces hold your users’ data. Prompts, tool inputs and outputs end up in them. Redact what you must (a hook, level 6), and treat the trace store like a database with personal data in it.
  • Cost is a metric too. Record tokens per span. An agent that passes every case with twice the calls is a regression your pass rate won’t show.