For the complete documentation index, see llms.txt. This page is also available as Markdown.

Evaluation spaces

Evaluation spaces — saved configurations you re-run as conditions change.

A space is a saved evaluation configuration. Bundle a model selection, a dataset, and a scoring config. Re-run the space whenever you want to evaluate the same configuration against fresh data, a fresh prompt, or a new candidate model.

Why spaces matter

Without spaces, every evaluation is a one-off — its config lives in someone's head or a runbook. Spaces formalize the configuration so:

  • Anyone on the team can re-run with one click

  • Run history is visible in one place

  • Score-over-time charts make trends obvious

Public vs private spaces

  • Public spaces — curated by LayerLens and partners; visible to everyone

  • Private spaces — scoped to your org

Both share the same shape.

Space contents

  • Models — one or more

  • Dataset / benchmark — one

  • Scoring config — scorers + judges

  • Optional schedule — re-run cadence

  • Run history — every run's results, with timestamps

Cloning

You can clone a public space into a private space, then customize. Useful when a public space matches your shape and you want to swap in your own dataset or judges.

When NOT to use a space

  • You're exploring — one-off evaluations are simpler

  • The configuration changes every run — a space is for stable configs

Where to next

Last updated

Was this helpful?