Manufacturing and industrial evaluation patterns
Manufacturing evaluation patterns — primitive ratio, span rules, GEPA labeling, cadence.
Last updated
Was this helpful?
Manufacturing evaluation patterns — primitive ratio, span rules, GEPA labeling, cadence.
Recommended primitive ratio for manufacturing AI: 65/25/10 — deterministic rules, code assertions, judges. Process safety, regulatory specs, and sanctions screening are rule-bound; defect taxonomies are taxonomy-coded. Judges are reserved for operator-readability and explanation quality.
Deterministic rules
65%
Sanctions screening, spec match, safety SOP citation, manual-version match
Code assertions
25%
Per-class P/R, drift, false-pass floors, country-of-origin
LLM judges
10%
Operator-readability, explanation quality, faithfulness to manual
Predictive maintenance — RUL calibration
Code assertion
Predicted vs. actual time-to-failure
Predictive maintenance — critical-asset escalation
Rule
Always-escalate for high-criticality
Predictive maintenance — citation
Judge
Recommendation grounded in failure mode
Quality — per-class P/R
Code assertion
Tracked separately per defect class
Quality — critical defect
Rule
Stop-and-hold for safety-critical
Quality — drift
Code assertion
Daily vs. golden image set
Supply chain — sanctions list
Rule
Every supplier OFAC/BIS/State checked
Supply chain — spec match
Rule
Substitution meets regulatory spec
Supply chain — country-of-origin
Code assertion
Verified database match
Field service — manual-version
Rule
Cited manual applies to machine
Field service — safety procedure
Rule
LOTO/PPE requirements present
Field service — out-of-scope
Rule
Refuse outside documented scope
Operator copilot — advisory only
Rule
AI never auto-applies set points
Operator copilot — handover completeness
Code scorer
Required fields present
Operator copilot — plain language
Code scorer
Reading level for operators
asset.lookup — machine ID, model, criticality classification
manual.cite — version, section, applicability check
sanctions.check — supplier-list match
spec.match — regulatory-spec verification
defect.classify — per-class scores
recommendation.advisory — flag confirming advisory-only status
For each manufacturing judge, label ≥ 50 paired examples:
Operator-readability: Floor ops + safety team labels
Field-service faithfulness: Senior techs + OEM trainers label, cover ambiguous troubleshoots
Recommendation grounding: Reliability engineers label against failure-mode database
Run judge optimization with budget="medium". Hold out 20%.
Per-deploy
CI gate runs scenario suite (rules + assertions)
Per-shift
Critical-asset alarm summary; quality drift report
Daily
Sample production traces; full evaluation
Weekly
Sanctions list refresh; spec-match audit
Monthly
GEPA re-optimization; manual-version refresh
Quarterly
Full process-safety audit; OSHA / EPA RMP prep
Annually
Recall-readiness drill; OEM-bulletin refresh
Last updated
Was this helpful?
Was this helpful?