For the complete documentation index, see llms.txt. This page is also available as Markdown.

Recipe: detect tool-call regressions in agents

Recipe — detect tool-call regressions. Pointer to the SDK code-review cowork sample.

Canonical samples:

Pattern

Capture full agent traces including tool-call spans. Build a judge whose rubric forbids specific tool calls in specific contexts (e.g., destructive APIs only when input scope says allowed). Run via client.trace_evaluations.create(trace_id=, judge_id=). Compare to baseline via samples/core/compare_evaluations.py.

See also

Last updated

Was this helpful?