Observe (workflow)
Stage 2 — Observe. See real production behavior across surfaces.
Last updated
Was this helpful?
Stage 2 — Observe. See real production behavior across surfaces.
Stage 2. With traces flowing in, observe what your AI is actually doing in production.
"What is my AI doing — and where is it surprising me?"
Browse the trace list. Filter by tags, by time range, by span errors. Look for outliers.
Drill into a sampled subset. Pick 20-30 traces from different time slices; read them.
Tag patterns. When you find an interesting trace shape (a known good, a known bad, an edge case), tag it. Tags become the foundation of your test set.
Notice latency and cost outliers. Span-level latency and cost surface in the trace inspector.
A working sense of what production looks like
A small tagged set of "interesting" traces — your starter trace set for evaluation
Hypotheses about where quality might drift
Don't define gates yet — you don't have enough signal
Don't optimize prompts yet — you'll be optimizing on hunches
Last updated
Was this helpful?
Was this helpful?