> For the complete documentation index, see [llms.txt](https://docs.layerlens.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.layerlens.ai/more-in-this-section-6/score-code.md).

# Recipe: evaluate code-generation models

Recipe — evaluate code-generation models. Use a code benchmark from the public catalog.

**When to use:** model returns code; correctness is measured by passing tests.

Evaluate against the relevant code benchmark from the [public catalog](/5.-select-pick-the-model/benchmarks-catalog.md):

```python
from layerlens import Stratix
client = Stratix()

# Pick a code benchmark from the public catalog
benchmark = client.benchmarks.get_by_key("humaneval") # or mbpp, swe-bench, etc.
model = client.models.get_by_key("openai/gpt-4o")

evaluation = client.evaluations.create(model=model, benchmark=benchmark)
evaluation = client.evaluations.wait_for_completion(evaluation, timeout_seconds=1800)

print(f"Accuracy: {evaluation.accuracy}")
```

For your own code-evaluation set with custom test harnesses, upload a JSONL of `{"input": "...", "truth": "..."}` rows via `client.benchmarks.create_custom(...)` — see [SDK reference: models-benchmarks](/more-in-this-section-9/models-benchmarks.md).

For domain-specific code review (security, style, correctness against your codebase conventions), use the [Cowork code-review pattern](https://github.com/layerlens/stratix-python/blob/main/samples/cowork/code_review.py) — Instrumentor uploads code traces, Reviewer evaluates with focused judges.

## See also

* [Concept: Models and benchmarks](/5.-select-pick-the-model/models-and-benchmarks.md)
* [Stratix Public — Benchmarks catalog](/5.-select-pick-the-model/benchmarks-catalog.md)
* [SDK sample: code-review cowork](https://github.com/layerlens/stratix-python/blob/main/samples/cowork/code_review.py)
