Public evaluation spaces
Public evaluation spaces — curated bundles of model + dataset + scoring config, browsable without signup.
Last updated
Was this helpful?
Public evaluation spaces — curated bundles of model + dataset + scoring config, browsable without signup.
An evaluation space is a workspace bundling a model selection, a dataset/benchmark selection, and a scoring config. Public spaces are curated by LayerLens and partners — browse them to see how seasoned teams structure evaluations.
Model selection — one or more models being evaluated
Dataset / benchmark selection — what's being evaluated against
Scoring config — which scorers and judges
Latest run results — top-line scores and detail
Description and methodology notes — context for why this space exists
Read for inspiration. A well-structured space is a recipe for "how to evaluate X."
Replicate. Use the same models, the same benchmarks, the same scoring on your own dataset. Premium evaluations let you instantiate a similar configuration.
Compare. Run the same evaluation against your data and see how your results compare to the public space.
Search by name or description
Filter by capability — code, reasoning, multilingual, etc.
Filter by industry tag
Sort by recency or by activity
Premium adds private spaces that are scoped to your org. The shape is identical — model + dataset + scoring + results — but visibility is private.
Last updated
Was this helpful?
Was this helpful?