For the complete documentation index, see llms.txt. This page is also available as Markdown.

Compare models

Compare any two models head-to-head across any shared benchmark — Stratix Public's flagship feature.

The Compare models feature is Stratix Public's flagship: pick any two models, see them head-to-head across every benchmark either has been evaluated against.

What you can do

  • Pick any two models from the catalog

  • See a side-by-side score table across all shared benchmarks

  • Drill into a single benchmark for a deeper comparison

  • See where each model wins, ties, and loses

  • Bookmark the comparison URL — it's stable

How it's structured

The compare-models page has three regions:

  1. Selector — search and pick model A and model B

  2. Score table — one row per benchmark, model-A score, model-B score, winner

  3. Detail view — click any benchmark to expand: raw scores, sample inputs, sample outputs

How comparisons render

For each benchmark:

  • Stronger model gets a green check + score

  • Weaker model gets the score plain

  • Tie is shown gray

  • The gap (numeric and percentage) is shown

When the gap is small, the page shows a confidence band — was the difference statistically meaningful given the benchmark sample size?

Three or more models

The Premium Evaluations surface supports comparing many models on your own dataset; the public Compare-models page is intentionally pairwise (this matches how most users decide).

Common patterns

  • Frontier vs cost-optimized — the same provider's flagship vs a smaller, cheaper sibling

  • Provider against provider — OpenAI's flagship vs Anthropic's vs Google's

  • Old vs new generation — when a new model drops, immediately compare against its predecessor

Where to next

Last updated

Was this helpful?