Benchmarks catalog
Stratix Public Benchmarks catalog — 52+ benchmarks with methodology and per-model scores.
Last updated
Was this helpful?
Stratix Public Benchmarks catalog — 52+ benchmarks with methodology and per-model scores.
The Stratix Public Benchmarks catalog is the broadest open-browsable collection of LLM benchmarks. 52+ benchmarks, each with documented methodology, sample tasks, and per-model scores.
Description — what the benchmark measures
Methodology — how scoring works, sample size, harness notes
Sample tasks — representative inputs from the benchmark
Per-model scores — every model evaluated against this benchmark
Top performers — leaderboard for this specific benchmark
Score-history-over-time — how the frontier moved on this benchmark
Capability — reasoning, code, math, multilingual, vision, multi-turn
Difficulty — easy / medium / hard
Sample size — small / medium / large
License — open / restricted
Alphabetical
By difficulty
By number of models evaluated
By recency
Each benchmark has a dedicated page with:
Hero card (name, capability, sample size)
Full methodology notes
Sample tasks (a representative subset)
Score table with every model's result
Top-N leaderboard
Score-history chart
Linked public evaluations
The hardest part of using benchmarks is picking the right ones. Some heuristics:
Don't pick more than 3-5. More benchmarks doesn't mean more signal.
Match capability to task. Code task → HumanEval/MBPP. Math task → MATH/GSM8K. General reasoning → MMLU/ARC.
Validate that the benchmark's distribution matches yours. A model that scores high on a benchmark drawn from a totally different domain doesn't help you.
Each quarterly research report documents methodology changes for the quarter — new benchmarks added, scoring changes, harness updates. This is the most rigorous public discussion of benchmark methodology you'll find.
The Public catalog scores models against an open library of benchmarks. If your task is domain-specific — pricing in your tariff, citing your jurisdiction, answering against your documents — public benchmarks narrow the field but won't decide.
Stratix Premium lets you author custom benchmarks from your own data and run the same candidate models against them. Custom benchmarks live in your workspace, are versioned, and rerun on demand whenever a new model lands in the catalog. See Stratix Premium → Benchmarks for the custom-benchmark workflow.
Last updated
Was this helpful?
Was this helpful?