For the complete documentation index, see llms.txt. This page is also available as Markdown.

Recipe: catch hallucinations

Recipe — catch hallucinations with a faithfulness judge whose rubric handles refusals.

When to use: outputs make factual claims; you need a guard against invented facts.

A faithfulness judge whose rubric explicitly handles refusals ("I don't know" should not score as a hallucination) is the canonical Stratix shape:

from layerlens import Stratix
client = Stratix()

faithfulness = client.judges.create(
 name="faithfulness-with-refusal-handling",
 evaluation_goal="""Score 0-1 for whether the OUTPUT is faithful to provided source CONTEXT.

Score 1 if:
- Every factual claim in OUTPUT is supported by CONTEXT
- OUTPUT explicitly refuses to answer ('I don't know', 'this isn't covered') when CONTEXT lacks the information

Score 0 if:
- OUTPUT makes any factual claim not supported by CONTEXT
- OUTPUT confidently answers when CONTEXT lacks the information (this is the worst failure mode)

Penalize confident-sounding hallucinations the most.""",
)

# Apply against traces that capture context + output
trace_eval = client.trace_evaluations.create(trace_id=trace.id, judge_id=faithfulness.id)
result = client.trace_evaluations.wait_for_completion(trace_eval.id)

GEPA-optimize against ≥50 labeled examples to push agreement-with-humans above 90%. See Bootstrap a judge before GEPA.

See also

Last updated

Was this helpful?