For the complete documentation index, see llms.txt. This page is also available as Markdown.

Evaluations

The Stratix Premium Evaluations surface — create, run, browse, and compare private evaluations.

Available in Stratix Premium. This surface is part of the logged-in workspace at stratix.layerlens.ai. Stratix Public users can browse the catalog but cannot use this feature.

The Evaluations page is where you run, browse, and compare private evaluation runs. It's the most-used surface in Premium.

What you can do

  • Create a new evaluation — pick a model, dataset, and scoring config

  • Browse past runs — filter by model, benchmark, scorer, judge, date

  • Re-run an evaluation as configuration changes

  • Compare two or more runs side-by-side

  • Drill into a single run — per-row results, score distribution, latency, cost

  • Export results to CSV/JSON for downstream analysis

Creating an evaluation

The new-evaluation flow has 5 steps:

  1. Pick model(s) — one or more models, including your BYOK custom models

  2. Pick dataset / benchmark — upload, select from public benchmarks, or pick your private dataset

  3. Pick scoring — code graders, judges, or both

  4. Preview cost — Stratix shows worst-case ECU consumption before you run

  5. Run — queue and watch results stream in

Browsing past runs

Filters in the sidebar:

  • Date range

  • Model

  • Benchmark / dataset

  • Scorer / judge

  • Status (queued, running, completed, failed)

  • Tags

Each row shows: name, model, benchmark, top-line score, status, ECU consumed, "compare" button.

Comparing runs

Select 2+ rows and click Compare. The comparison view shows:

  • Side-by-side score tables

  • Per-row deltas (where they ran on the same dataset)

  • Cost and latency comparison

Run a comparison against the public catalog

The compare-models view in Premium is identical in shape to Stratix Public's compare-models — you can compare your BYOK custom model against any public model.

Where to next

Last updated

Was this helpful?