For the complete documentation index, see llms.txt. This page is also available as Markdown.

Retail and e-commerce scenarios

Retail and e-commerce scenarios — product Q&A, recommendations, pricing, customer service, returns/fraud. Industry patterns, not customer case studies.

Five canonical scenarios where retailers and marketplaces use Stratix evaluation to gate AI features into production.

1. Product discovery and Q&A

Setup: Customer asks "Will this jacket fit a 6'2" 220lb frame?" "Is this dishwasher compatible with my hardwired connection?" AI synthesizes the product spec sheet, prior reviews, manufacturer Q&A; produces an answer with source citations.

Why it's hard:

  • Product attributes may be missing or inconsistent across the catalog

  • Reviews can contradict spec sheets

  • Customers act on the answer — wrong fit info drives returns; wrong compatibility drives serious damage

Stakes:

  • Returns directly hit margin; some categories see 30%+ return rates

  • Compatibility errors on regulated/installed products (appliances, automotive parts) can trigger product-liability exposure

  • FTC and state AGs flag misleading product claims even when AI-generated

Stratix evaluation criteria:

Aspect
What you measure
Pattern

Spec grounding

Output cites the spec sheet

Citation existence

Review-spec conflict

Output flags conflicts vs. picking arbitrarily

Code assertion

Compatibility hard rule

High-risk categories require explicit compatibility-check disclaimer

Hard rule

Hallucination rate

Confident wrong claims about specs

Faithfulness judge

Reading level

Comparable to category baseline

Code scorer

2. Personalized recommendations

Setup: Recommendation engine surfaces products. AI-augmented version uses LLM to expand "similar products," generate copy, or rerank.

Why it's hard:

  • Stale catalog — recommendation surfaces a discontinued SKU

  • Sensitive-category guardrails — no recommending alcohol/firearms/adult on a kid's account

  • Personalization fairness — same product should not vary in price/availability based on protected-class proxies

Stakes:

  • Bait-and-switch perception when out-of-stock items dominate recommendations

  • FTC scrutiny on dark patterns and personalized pricing

  • State AG actions on sensitive categories (alcohol, tobacco, firearms) reaching wrong audience

Stratix evaluation criteria:

  • Catalog freshness rule — Hard rule: every recommended SKU must be in-stock and live in the catalog as of trace time

  • Sensitive-category gate — Hard rule by account / age band

  • Personalization fairness — Group fairness on availability and surfaced price

  • Diversity scorer — Recommendation-set diversity to avoid filter-bubble effects

  • Conversion-correlation watch — Soft drift signal; recommend regenerating evals if conversion drops

3. Customer service chatbot

Setup: Order status, returns, exchanges, escalation. Order data + policy data + customer history.

Why it's hard:

  • Authoritative policy must be reflected (return windows, restocking fees)

  • Refund-promise accuracy — bot promises a refund the policy doesn't allow

  • Multi-channel consistency — same answer in chat, email, voice

Stakes:

  • Refund-promise mismatch → regulatory and reputational exposure

  • CCPA / GDPR data subject requests must route correctly

  • FTC focuses heavily on dark-pattern returns/cancellation flows

Stratix evaluation criteria:

  • Policy citation rule — Every refund/return decision cites the relevant policy section

  • Refund-promise gate — Hard rule: bot cannot promise refunds outside policy parameters

  • DSR routing — Privacy requests recognized and routed to the right intake

  • Tone and de-escalation — Judge for customer-experience quality

  • Channel parity — Same query produces same answer across chat/email/voice

4. Dynamic pricing and promotion

Setup: AI surfaces pricing or promotion suggestions. Human merchandiser approves; AI does not autonomously change prices.

Why it's hard:

  • Personalized-pricing legal exposure — if it correlates with protected class, it's discriminatory

  • Bait pricing — promotion that's never actually realized

  • Cross-region price-parity expectations

Stakes:

  • FTC and state AG actions on personalized pricing have multiplied in 2025-2026

  • Class-action exposure under state UDAP statutes

  • Marketplace platform-policy violations carry delisting risk

Stratix evaluation criteria:

  • Personalized-price fairness — Group fairness on offered price across protected proxies

  • Realizability rule — Hard rule: any "promotion" must be actually achievable, not bait

  • Cross-region parity — Code scorer flagging unjustified regional spread

  • Audit trail — Every price suggestion logged with inputs, model version, approver

5. Returns and fraud detection

Setup: AI scores returns for fraud markers (wardrobing, empty-box returns, serial returners, organized retail crime patterns). Outputs go to fraud team for action.

Why it's hard:

  • False positives wrongly accuse legitimate customers and drive churn

  • Disparate impact on legitimate but high-return categories (apparel, gifts)

  • Privacy boundaries on what the score can use as features

Stakes:

  • Retail loss to returns fraud reaches multi-billions annually across the industry

  • Wrongful denials drive class-action exposure

  • State consumer-protection AGs investigate algorithmic denial of returns

Stratix evaluation criteria:

  • Top-decile precision — Tracked against confirmed-fraud labels

  • Fairness scorer — Group-level denial rate parity

  • Reason-code coverage — Hard rule: every flagged return cites contributing factors

  • Customer-explanation rule — Hard rule: customer-facing denial reason is plain-language

  • Override rate watch — Drift signal if humans override AI more than threshold

See also

Last updated

Was this helpful?