For the complete documentation index, see llms.txt. This page is also available as Markdown.

Observe (workflow)

Stage 2 — Observe. See real production behavior across surfaces.

Stage 2. With traces flowing in, observe what your AI is actually doing in production.

The question this stage answers

"What is my AI doing — and where is it surprising me?"

What to do

  1. Browse the trace list. Filter by tags, by time range, by span errors. Look for outliers.

  2. Drill into a sampled subset. Pick 20-30 traces from different time slices; read them.

  3. Tag patterns. When you find an interesting trace shape (a known good, a known bad, an edge case), tag it. Tags become the foundation of your test set.

  4. Notice latency and cost outliers. Span-level latency and cost surface in the trace inspector.

What you'll have at the end

  • A working sense of what production looks like

  • A small tagged set of "interesting" traces — your starter trace set for evaluation

  • Hypotheses about where quality might drift

Don't yet

  • Don't define gates yet — you don't have enough signal

  • Don't optimize prompts yet — you'll be optimizing on hunches

Where to next

Last updated

Was this helpful?