For the complete documentation index, see llms.txt. This page is also available as Markdown.

Browse benchmarks

Browse 52+ public benchmarks with descriptions, methodology, and per-model scores.

The Stratix Public Benchmarks catalog is the broadest open-browsable collection of LLM benchmarks today. Each benchmark has metadata, methodology notes, and per-model scores.

Steps

1. Open the catalog

Go to stratix.layerlens.ai/benchmarks.

2. Filter

  • Capability — reasoning, code, math, multilingual, vision, multi-turn

  • Difficulty — easy / medium / hard

  • Sample size — small / medium / large

3. Open a benchmark

Click any benchmark card. The page shows:

  • Description and methodology

  • Sample tasks

  • Per-model scores

  • Top performers

  • Score-history-over-time chart

  • Public evaluations that ran this benchmark

4. Pick a model and see its score

From the benchmark page, click any model in the score table to jump to that model on this benchmark.

Verify

You should be able to find MMLU, HumanEval, GSM8K on the first page.

How to pick a benchmark for your use case

Anti-pattern: "everyone uses MMLU, so I'll use MMLU."

Instead:

  • Reasoning-heavy task? Look at GSM8K, MATH, ARC.

  • Code? HumanEval, MBPP, SWE-Bench.

  • Multi-turn dialog? MT-Bench.

  • Multilingual? MMLU translated, FLORES.

  • Tool use / agents? ToolBench, AgentBench.

  • Reading comprehension? SQuAD, DROP.

Where to next

Last updated

Was this helpful?