Energy and utilities evaluation patterns
Energy and utilities evaluation patterns — primitive ratio, span rules, GEPA labeling, cadence.
Last updated
Was this helpful?
Energy and utilities evaluation patterns — primitive ratio, span rules, GEPA labeling, cadence.
Recommended primitive ratio for energy AI: 70/20/10 — deterministic rules, code assertions, judges. Grid procedures, NERC CIP, OSHA / DOT safety rules, and ISO/RTO market rules drive heavy rule density. Judges are reserved for customer-facing copy quality.
Deterministic rules
70%
Switching procedures, NERC reliability, ISO market, medical-baseline, LOTO/PPE
Code assertions
20%
Risk calibration, freshness thresholds, settlement-data checks
LLM judges
10%
Customer-impact statement quality, multilingual parity, switching-step faithfulness
Grid ops — switching
Rule + Judge
Output matches approved procedures
Grid ops — N-1
Rule
Contingency obligations preserved
Grid ops — advisory only
Rule
AI never auto-dispatches
Predictive maintenance — calibration
Code assertion
Predicted vs. actual failures
Predictive maintenance — wildfire
Rule
High-fire-threat priority
Predictive maintenance — vegetation
Rule
State VM standards followed
Customer service — medical baseline
Rule
Always-route-to-human
Customer service — low-income program
Rule
Eligibility surfaced
Customer service — outage freshness
Rule
Data within threshold
DR / DERMS — ISO rule compliance
Rule
Match ISO/RTO market rules
DR / DERMS — settlement freshness
Code assertion
Meter data current
DR / DERMS — cyber gate
Rule
NERC CIP controls preserved
Field safety — energized work
Rule
LOTO/PPE requirements
Field safety — pipeline procedure
Rule
49 CFR cited
Field safety — manual version
Rule
Cited manual current
procedure.cite — switching/maintenance procedure version
nerc.check — applicable NERC standard
iso.market.rule — applicable ISO market rule
wildfire.zone — fire-threat district classification
medical.baseline — protected-customer routing
safety.disclaimer — required disclaimer presence
For each energy judge, label ≥ 50 paired examples:
Switching-step faithfulness: Senior dispatchers + reliability engineers label
Customer-impact messages: Comms team + customer advocates label
Field-safety faithfulness: Senior linemen / pipeline-qualified leads label
Multilingual: Native-speaker QA per language
Run judge optimization with budget="medium". For safety-critical judges, run with budget="heavy" and hold out 30%.
Per-deploy
CI gate runs scenario suite (rules + assertions)
Per-shift
Critical-asset alarm summary; advisory-only audit
Continuous
NERC alerts; outage-data freshness
Daily
Sample production traces; full evaluation
Weekly
Wildfire-zone audit (Cal Fire / state); vegetation status
Monthly
GEPA re-optimization on judges; switching-procedure version refresh
Quarterly
NERC CIP audit prep; ISO/RTO market-rule refresh
Annually
OSHA / PHMSA reportable-event prep; full safety audit
Last updated
Was this helpful?
Was this helpful?