Build a trace set for agentic evals
Guide — build a trace set for agentic evaluation.
Last updated
Was this helpful?
Guide — build a trace set for agentic evaluation.
Identify representative inputs (happy paths, edge cases, known-bad).
Run your agent against each; capture as a trace.
Tag and group them into a dataset (a named, versioned trace set).
Use the dataset as input to agentic evaluation.
You don't need production traffic to build a trace set. From Catalog → Datasets → New Dataset → Generate synthetic traces, produce a realistic multi-agent dataset from a built-in industry scenario (or by varying a few real traces) in minutes, then evaluate it like any other dataset. See Synthetic data.
Last updated
Was this helpful?
Was this helpful?