Evaluation spaces
Evaluation spaces — saved configurations you re-run as conditions change.
Last updated
Was this helpful?
Evaluation spaces — saved configurations you re-run as conditions change.
A space is a saved evaluation configuration. Bundle a model selection, a dataset, and a scoring config. Re-run the space whenever you want to evaluate the same configuration against fresh data, a fresh prompt, or a new candidate model.
Without spaces, every evaluation is a one-off — its config lives in someone's head or a runbook. Spaces formalize the configuration so:
Anyone on the team can re-run with one click
Run history is visible in one place
Score-over-time charts make trends obvious
Public spaces — curated by LayerLens and partners; visible to everyone
Private spaces — scoped to your org
Both share the same shape.
Models — one or more
Dataset / benchmark — one
Scoring config — scorers + judges
Optional schedule — re-run cadence
Run history — every run's results, with timestamps
You can clone a public space into a private space, then customize. Useful when a public space matches your shape and you want to swap in your own dataset or judges.
You're exploring — one-off evaluations are simpler
The configuration changes every run — a space is for stable configs
Last updated
Was this helpful?
Was this helpful?