For the complete documentation index, see llms.txt. This page is also available as Markdown.

Recipe: evaluate code-generation models

Recipe — evaluate code-generation models. Use a code benchmark from the public catalog.

When to use: model returns code; correctness is measured by passing tests.

Evaluate against the relevant code benchmark from the public catalog:

from layerlens import Stratix
client = Stratix()

# Pick a code benchmark from the public catalog
benchmark = client.benchmarks.get_by_key("humaneval") # or mbpp, swe-bench, etc.
model = client.models.get_by_key("openai/gpt-4o")

evaluation = client.evaluations.create(model=model, benchmark=benchmark)
evaluation = client.evaluations.wait_for_completion(evaluation, timeout_seconds=1800)

print(f"Accuracy: {evaluation.accuracy}")

For your own code-evaluation set with custom test harnesses, upload a JSONL of {"input": "...", "truth": "..."} rows via client.benchmarks.create_custom(...) — see SDK reference: models-benchmarks.

For domain-specific code review (security, style, correctness against your codebase conventions), use the Cowork code-review pattern — Instrumentor uploads code traces, Reviewer evaluates with focused judges.

See also

Last updated

Was this helpful?