> For the complete documentation index, see [llms.txt](https://docs.layerlens.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.layerlens.ai/more-in-this-section-6/catch-hallucinations.md).

# Recipe: catch hallucinations

Recipe — catch hallucinations with a faithfulness judge whose rubric handles refusals.

**When to use:** outputs make factual claims; you need a guard against invented facts.

A faithfulness judge whose rubric explicitly handles refusals ("I don't know" should not score as a hallucination) is the canonical Stratix shape:

```python
from layerlens import Stratix
client = Stratix()

faithfulness = client.judges.create(
 name="faithfulness-with-refusal-handling",
 evaluation_goal="""Score 0-1 for whether the OUTPUT is faithful to provided source CONTEXT.

Score 1 if:
- Every factual claim in OUTPUT is supported by CONTEXT
- OUTPUT explicitly refuses to answer ('I don't know', 'this isn't covered') when CONTEXT lacks the information

Score 0 if:
- OUTPUT makes any factual claim not supported by CONTEXT
- OUTPUT confidently answers when CONTEXT lacks the information (this is the worst failure mode)

Penalize confident-sounding hallucinations the most.""",
)

# Apply against traces that capture context + output
trace_eval = client.trace_evaluations.create(trace_id=trace.id, judge_id=faithfulness.id)
result = client.trace_evaluations.wait_for_completion(trace_eval.id)
```

GEPA-optimize against ≥50 labeled examples to push agreement-with-humans above 90%. See [Bootstrap a judge before GEPA](/more-in-this-section-6/bootstrap-judges.md).

## See also

* [Concept: Judges](/8.-evaluate-score-the-outputs/judges-1.md)
* [Use case: RAG evaluation](/4.1-general-use-cases/rag-evaluation.md)
* [SDK sample: rag\_assessment](https://github.com/layerlens/stratix-python/blob/main/samples/cowork/rag_assessment.py)
