Retail and e-commerce evaluation patterns
Retail evaluation patterns — primitive ratio, span rules, GEPA labeling, cadence.
Last updated
Was this helpful?
Retail evaluation patterns — primitive ratio, span rules, GEPA labeling, cadence.
Recommended primitive ratio for retail AI: 55/30/15 — deterministic rules, code assertions, judges. Retail has more open-text customer interaction than financial services or insurance, so judges carry slightly more weight.
Deterministic rules
55%
Catalog freshness, policy match, sensitive-category gates, refund-promise limits
Code assertions
30%
Group fairness, diversity, channel parity, top-decile precision
LLM judges
15%
Tone/de-escalation, customer-experience quality, hallucination detection
Product Q&A — spec citation
Rule
Output cites spec sheet
Product Q&A — hallucination
Judge
Faithfulness to spec
Product Q&A — compatibility disclaimer
Rule
Required for high-risk categories
Recommendations — catalog freshness
Rule
All SKUs in-stock and live
Recommendations — sensitive category
Rule
Age/account-band gating
Recommendations — fairness
Code scorer
Group-level price/availability parity
Recommendations — diversity
Code scorer
Set diversity threshold
Customer service — policy citation
Rule
Refund decisions cite policy section
Customer service — refund promise
Rule
No promises outside policy
Customer service — DSR routing
Rule
Privacy requests routed correctly
Customer service — tone
Judge
De-escalation quality
Pricing — personalized fairness
Code scorer
Group-level price parity
Pricing — realizability
Rule
Promotions must be achievable
Returns fraud — top-decile precision
Code assertion
Match confirmed-fraud labels
Returns fraud — denial reason
Rule
Plain-language reason mandatory
catalog.lookup — SKU live, in-stock, in-region
policy.cite — return/refund/promotion policy cited
personalization.score — feature inputs and bias check
sensitive.gate — age/category gating decision
fraud.score — score, top contributing features
For each retail judge, label ≥ 50 paired examples:
Customer-service tone: Label by senior CX agents; cover de-escalations, simple resolutions, hard escalations
Product Q&A faithfulness: Spec-grounded ground-truth; cover ambiguous review-vs-spec cases
Returns fraud: Confirmed-fraud + non-fraud controls; 1:5 ratio works for most retailers
Recommendation diversity: Buyer-team labeled "good" vs "filter-bubble" recommendation sets
Run judge optimization with budget="medium". Hold out 20%.
Per-deploy
CI gate runs scenario suite (rules + assertions)
Hourly
Catalog-freshness rule against live SKU feed
Daily
Sample production traces across all five scenarios; full scoring
Weekly
Personalized-pricing fairness audit; returns-fraud override-rate watch
Monthly
GEPA re-optimization on judge drift
Quarterly
Full bias re-audit; sensitive-category gate review
Last updated
Was this helpful?
Was this helpful?