From Arize
Migrate from Arize AI to Stratix — concept mapping, what to port, what to re-author, and a phased cutover plan.
Arize and Stratix overlap on trace observability + LLM evaluation but Arize is broader (covering classical ML observability) while Stratix is LLM-evaluation-first. Migration usually means keeping Arize for classical-ML observability if you have it, and moving LLM-specific evaluation workflows to Stratix. Most teams complete LLM-cutover in 2–3 weeks running both systems in parallel.
Concept mapping
Space
Project (within Organization)
Both scope evaluation work
Inference
Trace row
Arize stores per-prediction; Stratix's trace can hold many predictions plus tool calls and retrieval steps
Span (Phoenix trace)
Span (Stratix trace)
Both follow OpenInference / OpenTelemetry conventions; field names differ slightly
Hierarchical trace tree
Stratix trace tree
One-to-one mapping; Stratix natively models multi-agent handoffs as spans
Dataset
Custom benchmark
Stratix benchmarks are versioned and rerunnable
Evaluator
Choose: Scorer (LLM-prompt, reusable across benchmarks) or Judge (LLM-rubric, versioned, GEPA-tunable) or Code grader (deterministic check in the evaluation runtime)
Drift monitor
Continuous trace evaluation + threshold alert
Embeddings monitor
Stratix continuous evaluation pulls trace samples and runs scorers/judges; embedding-drift specifically is a custom code grader pattern today
Performance metrics dashboard
Stratix Home dashboard + per-evaluation metrics
Annotation queue
Labeling workflow on trace evaluation → feeds GEPA
Trace search
Stratix trace search with structured filters
What does NOT map cleanly
Classical-ML observability (regression, tabular classification, recommender ranking metrics). Stratix is not a classical-ML observability platform. If you use Arize for both ML and LLM monitoring, keep Arize for ML and move LLM-specific workflows to Stratix.
Embedding-drift monitors require a custom code-grader implementation in Stratix today; the canonical Arize embedding-drift dashboards don't have a direct equivalent.
Arize Copilot doesn't have a direct equivalent; the closest is the Stratix In-app Assistant.
Migration steps (phased cutover)
Phase 1 — Inventory (Day 1)
List every Arize space, dataset, evaluator, and drift monitor that touches LLM workloads.
For each evaluator, classify: LLM-judge, LLM-prompt scorer, code grader, embedding-drift (special-case).
Identify drift monitors with active paging — these need the cleanest cutover.
Phase 2 — Port traces (Days 2–4)
Trace export shape. Arize trace exports follow OpenInference conventions (spans with attributes.llm.input_messages, attributes.llm.output_messages, attributes.tool.name, etc.). Stratix's trace schema is also OpenInference-compatible — the mapping is mostly 1:1 with a few field-name normalizations.
Bulk export. Pull traces from Arize using Phoenix's
client.get_evaluations(...)andclient.get_spans_dataframe(...). Or, if you're using the Phoenix OTLP endpoint, point your OTLP exporter at Stratix directly going forward.Normalize field names. A small Python script reshapes OpenInference attribute names to Stratix's expected shape — see the Stratix trace schema. The recipe at Cookbook: backfill traces from logs covers the normalization pattern.
Upload via SDK in batches of ≤ 50 MB JSONL:
Phase 3 — Re-author evaluators (Days 4–10)
For Phoenix LLM evaluators:
Identify the rubric prompt
Decide: per-row inside an evaluation → Scorer; on traces directly → Judge
Author via SDK (
client.scorers.create(...)orclient.judges.create(...)) or the dashboardIf you have ≥ 30 labeled examples, run GEPA optimization on the judge to push agreement-with-humans up
For deterministic Arize evaluators (regex match, numeric thresholds, JSON schema validators):
Author as code graders in the evaluation runtime
For drift monitors:
Replace with continuous trace evaluation — schedule the corresponding judge/scorer to run on production trace samples at a cadence (hourly, daily) with threshold alerts on the resulting score distribution
Phase 4 — Dual-run (Days 11–18)
Send the same traces to both Arize and Stratix for one full cycle.
Compare per-trace scores; investigate divergence > 5%.
Adjust Stratix judges / scorers where divergence is rubric-quality, not data-quality.
Phase 5 — Cut over (Day 19+)
Switch your OTLP / trace-ingestion sink from Arize to Stratix (or run both if you want classical-ML coverage in Arize alongside LLM in Stratix).
Move alert/paging configurations.
Archive the LLM-specific portions of your Arize space; document the cut date for audit.
Common cutover gotchas
Trace-id collisions. If you re-import historical Arize traces and also stream new traces, conflicts can occur. Use a namespace prefix (
arize-) on imported trace IDs to avoid clashes.Field-name normalization. Arize sometimes uses Phoenix-specific attribute names; Stratix uses the OpenInference baseline. A small mapping table handles the few divergent fields.
Embedding-drift dashboards don't auto-port. Plan to author this as a custom code grader if you depend on it.
See also
Last updated
Was this helpful?