For the complete documentation index, see llms.txt. This page is also available as Markdown.

Benchmarks catalog

Stratix Public Benchmarks catalog — 52+ benchmarks with methodology and per-model scores.

The Stratix Public Benchmarks catalog is the broadest open-browsable collection of LLM benchmarks. 52+ benchmarks, each with documented methodology, sample tasks, and per-model scores.

URL: stratix.layerlens.ai/benchmarks

What you can see for each benchmark

  • Description — what the benchmark measures

  • Methodology — how scoring works, sample size, harness notes

  • Sample tasks — representative inputs from the benchmark

  • Per-model scores — every model evaluated against this benchmark

  • Top performers — leaderboard for this specific benchmark

  • Score-history-over-time — how the frontier moved on this benchmark

Filtering

  • Capability — reasoning, code, math, multilingual, vision, multi-turn

  • Difficulty — easy / medium / hard

  • Sample size — small / medium / large

  • License — open / restricted

Sorting

  • Alphabetical

  • By difficulty

  • By number of models evaluated

  • By recency

Per-benchmark page

Each benchmark has a dedicated page with:

  • Hero card (name, capability, sample size)

  • Full methodology notes

  • Sample tasks (a representative subset)

  • Score table with every model's result

  • Top-N leaderboard

  • Score-history chart

  • Linked public evaluations

Picking a benchmark

The hardest part of using benchmarks is picking the right ones. Some heuristics:

  • Don't pick more than 3-5. More benchmarks doesn't mean more signal.

  • Match capability to task. Code task → HumanEval/MBPP. Math task → MATH/GSM8K. General reasoning → MMLU/ARC.

  • Validate that the benchmark's distribution matches yours. A model that scores high on a benchmark drawn from a totally different domain doesn't help you.

Quarterly methodology updates

Each quarterly research report documents methodology changes for the quarter — new benchmarks added, scoring changes, harness updates. This is the most rigorous public discussion of benchmark methodology you'll find.

Want to run your own benchmarks?

The Public catalog scores models against an open library of benchmarks. If your task is domain-specific — pricing in your tariff, citing your jurisdiction, answering against your documents — public benchmarks narrow the field but won't decide.

Stratix Premium lets you author custom benchmarks from your own data and run the same candidate models against them. Custom benchmarks live in your workspace, are versioned, and rerun on demand whenever a new model lands in the catalog. See Stratix Premium → Benchmarks for the custom-benchmark workflow.

Where to next

Last updated

Was this helpful?