Browse benchmarks
Browse 52+ public benchmarks with descriptions, methodology, and per-model scores.
Last updated
Was this helpful?
Browse 52+ public benchmarks with descriptions, methodology, and per-model scores.
The Stratix Public Benchmarks catalog is the broadest open-browsable collection of LLM benchmarks today. Each benchmark has metadata, methodology notes, and per-model scores.
Capability — reasoning, code, math, multilingual, vision, multi-turn
Difficulty — easy / medium / hard
Sample size — small / medium / large
Click any benchmark card. The page shows:
Description and methodology
Sample tasks
Per-model scores
Top performers
Score-history-over-time chart
Public evaluations that ran this benchmark
From the benchmark page, click any model in the score table to jump to that model on this benchmark.
You should be able to find MMLU, HumanEval, GSM8K on the first page.
Anti-pattern: "everyone uses MMLU, so I'll use MMLU."
Instead:
Reasoning-heavy task? Look at GSM8K, MATH, ARC.
Code? HumanEval, MBPP, SWE-Bench.
Multi-turn dialog? MT-Bench.
Multilingual? MMLU translated, FLORES.
Tool use / agents? ToolBench, AgentBench.
Reading comprehension? SQuAD, DROP.
Last updated
Was this helpful?
Was this helpful?