For the complete documentation index, see llms.txt. This page is also available as Markdown.

Public evaluations

2,000+ public evaluations browsable on Stratix Public — full results, sample inputs, methodology.

A public evaluation is a complete (model, dataset, scoring config, results) bundle published openly on stratix.layerlens.ai. 2,000+ runs are browsable today.

URL: stratix.layerlens.ai/evaluations

What you can see for each evaluation

  • The model and benchmark involved

  • The scoring config (code graders and/or LLM judges)

  • Top-line score

  • Per-row results

  • Sample inputs and outputs (where licensing permits)

  • Score distribution

  • Run metadata (date, methodology version, contributor)

Filtering

  • Model

  • Benchmark

  • Capability

  • Recency

Sorting

  • Most recent

  • By model

  • By benchmark

  • By score

Per-evaluation page

Each evaluation page shows:

  • Hero card (model + benchmark + top-line score)

  • Methodology notes

  • Score breakdown

  • Sample rows

  • Linked compare-models pairs

Why public evaluations matter

  • Reproducibility. Every evaluation has a stable URL you can cite.

  • Methodology transparency. See exactly how scoring was done.

  • Trust. When a vendor claims "model X is best at Y," you can check Stratix's evaluation, not their marketing.

Contributing public evaluations

If you've run an evaluation on your own data and want to publish it publicly, see Sharing public results.

Where to next

Last updated

Was this helpful?