For the complete documentation index, see llms.txt. This page is also available as Markdown.

AI quality gates in CI/CD

Block prompt and model regressions in your merge pipeline — turn evaluations into a hard gate.

Code regressions are caught by tests. AI regressions slip through unless you treat the eval as a test. The CI/CD-quality-gates use case wires Stratix evaluations into your pipeline so a prompt change that drops accuracy can't merge.

The shape of the work

  1. Pick your gating evaluation. A small, fast subset of your benchmarks and judges. Goal: meaningful signal in under 5 minutes.

  2. Add the eval as a CI step. GitHub Actions, GitLab CI, Buildkite — call the SDK or CLI with the eval config.

  3. Compare against baseline. Stratix maintains an evaluation history. The CI step compares the new run to the most recent main-branch run.

  4. Fail the build on regression. If any tracked dimension drops by more than your tolerance, fail.

  5. Annotate the PR. Post the score table back to the PR for reviewer context.

Why it works on Stratix

  • Evaluation history is built-in. No need to roll your own diff infrastructure.

  • SDK and CLI both support CI. Pick whichever feels native.

  • Compare-models lets you stage upgrades. Switch the model in CI, verify scores, then promote.

  • GEPA-optimized judges keep gates honest. Out-of-the-box judges flake; tuned ones don't.

Tools you'll use

Outcomes you should see

You'll know this is working when:

  • Zero AI regressions reach production because the CI gate caught them.

  • CI gate run time stays under 5 minutes even as the eval grows.

  • Tolerance is calibrated — false-positive rate <5%, false-negative rate <2%.

  • Engineers cite the eval result in the PR description as a matter of habit.

Anti-patterns

  • 5-hour CI evals. If your team waits 5 hours for a merge gate, they'll ignore it. Pick a fast subset.

  • Whole-org tolerance. Different repos have different quality bars. Set per-repo or per-feature tolerances.

  • No alerting on baseline drift. If the baseline silently drops over 4 PRs, you never noticed because each PR was within tolerance. Watch the trend.

Where to next

Last updated

Was this helpful?