Skip to content
astorlm
Language: English
← Map

Level 6

Streaming and events

A model writes its reply a few tokens at a time. Streaming hands you each piece as it’s written, and the loop’s events tell the rest of your app what the agent is doing, while it does it.
1/37 Bandoneón folds:
  • user
  • assistant
  • tool_result
Showtime. The Oracle sings, the notes come down the highway, and the person on guitar plays them for the crowd. First, a plain request: no streaming.

EventBus

The problem

Calling a model and waiting for the whole answer works in a script. In front of a person it doesn’t: a long reply can take ten or twenty seconds, and during all of them the screen shows nothing. An agent makes it worse, because one request can run several turns and tools before the final answer exists.

And it isn’t only the person who wants to follow along. The interface wants the text, your logs want every step, billing wants the tokens of each turn. If each of those has to be written into the loop, the loop ends up knowing about every screen, file and database in your app.

The solution

Two pieces that go together. Streaming: ask the provider for the reply as a stream, and it sends each piece of text as soon as the model writes it. The total time is the same, but the first word shows up in a fraction of a second instead of at the end.

Events: the loop announces every step on a bus, and doesn’t care who is listening. Your app subscribes the parts that care: the screen listens for text, the log for everything, the meter for the end of each turn. Adding a listener doesn’t touch the loop. These are the events astorlm emits on a normal run:

  • turn_start

    A turn begins: the loop is about to call the model.

  • text_delta

    A piece of the reply arrived. Many per turn.

  • assistant_message

    The model’s message is complete, text and tool calls included.

  • tool_execution_start

    A tool is about to run, with its input.

  • tool_execution_end

    The tool finished: its output, whether it failed, how long it took.

  • turn_end

    The turn is over, with the tokens its model call used.

  • session_end

    The run is over: completed, aborted, or failed.

Tool calls stream from the provider too, a few characters of their arguments at a time. The loop puts them back together before it runs them, so on the bus a tool call shows up whole, in assistant_message.

The cast

Same cast as always, on a rock stage.

The Oracle the model
The singer, up on the riser. Its reply comes down the highway as notes, one per word.
The highway the stream
Each text_delta is a handful of notes. Without streaming, nothing comes down until the end.
The guitarist your app
Plays each note as it reaches the frets, and the lyrics appear for the crowd.
The cable the EventBus
Every event runs along it. Only the listeners that subscribed to that event light up.
The listeners agent.on(…)
The lyrics screen ('text'), a tape deck ('event') and a token meter (turn_end).
Astor the loop
Runs the loop as usual, and announces each step on the bus as he goes.
The stage the tools
The front edge, the lighting desk and the pyro box: crowd_mood, set_lights and pyro.

The code

With astorlm: the provider always streams, so there is nothing to switch on. Subscribe with agent.on() before you call run(): 'text' for each piece of text, 'tool-start' and 'tool-end' for tools, 'event' for everything. Each call returns a function that unsubscribes.

From scratch: the loop from level 2 with two changes. The request asks for stream: true and reads the reply as server-sent events, gluing the pieces back together; and a list of listeners gets every step through one emit().

import { OpenAIProvider, createLocalAgent, tool } from 'astorlm'
import { z } from 'zod'

// The stage's functions, wrapped as tools.
const crowdMood = tool({
  name: 'crowd_mood',
  description: 'Look at the crowd from the front of the stage: how they feel, and what they came to hear.',
  schema: z.object({}),
  execute: async () => stage.readCrowd(), // your code
})

const setLights = tool({
  name: 'set_lights',
  description: 'Set the color of every light on the stage.',
  schema: z.object({ color: z.enum(['amber', 'red', 'blue', 'white']) }),
  execute: async ({ color }) => stage.lights(color), // your code
})

const pyro = tool({
  name: 'pyro',
  description: 'Fire the pyrotechnics on both sides of the stage. Once per show.',
  schema: z.object({}),
  execute: async () => stage.firePyro(), // your code
})

const agent = await createLocalAgent({
  // Any OpenAI-compatible endpoint: OpenAI, Ollama, LM Studio, vLLM, a proxy…
  provider: new OpenAIProvider({
    baseURL: 'http://localhost:11434/v1', // e.g. Ollama's default address
    model: 'your-model', // e.g. 'llama3.1', 'gpt-4o-mini'
    apiKey: 'YOUR_API_KEY', // local servers usually ignore it
  }),
  tools: [crowdMood, setLights, pyro],
  maxTurns: 10,
})

// Streaming is always on: subscribe before run(). Each listener only hears what it asked for.
// The lyrics screen: every piece of text, the moment it arrives.
agent.on('text', (piece) => lyrics.append(piece))

// The tape deck: every event the loop emits, in order.
agent.on('event', (event) => tape.write(event))

// The meter: the tokens each turn reported.
agent.on('event', (event) => {
  if (event.type === 'turn_end' && event.usage) meter.add(event.usage.inputTokens + event.usage.outputTokens)
})

// Each on() returns a function that unsubscribes.
const stopTape = agent.on('tool-start', ({ name }) => console.log('→', name))

const last = await agent.run('Open the show: read the crowd, set the mood, and greet them.')
stopTape()
console.log(last.content) // the same text, whole, once the run is over

What to watch

  • Streaming doesn’t make the model faster. The reply takes as long as before; what changes is when the person sees the first word. Measure time to first token, not only total time.
  • Listeners watch; they don’t steer. Whatever a listener returns is ignored, and the loop doesn’t wait for it. To block a tool or change what the model reads, you need a hook: the next level.
  • Keep listeners quick. astorlm’s bus calls them one after the other, inside the loop. Slow work, like writing to a database, belongs in a queue the listener only pushes to.
  • A broken listener fails quietly. astorlm catches its error so the run goes on, which also means nobody hears about it. Log inside your listeners.
  • Pieces don’t respect words. A text_delta can end in the middle of a word or of a Markdown table. Render what you have so far, and redraw as more arrives.
  • Let the person stop it. Once they can watch the reply arrive, they will want to cut a wrong one short. Wire a stop button to agent.abort().