npx skills add ...
npx skills add arize-ai/arize-skills --skill arize-evaluator
Handles LLM-as-judge and code evaluator workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, code evaluator, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.
npx skills add arize-ai/arize-skills --skill arize-evaluator
SPACE—--spaceflags accept a space name (e.g.,my-workspace) or a base64 space ID (e.g.,U3BhY2U6...). Find yours withax spaces list.
This skill covers designing, creating, and running evaluators on Arize — both LLM-as-judge (template) evaluators and code evaluators (deterministic, no LLM required). An evaluator defines the judge; a task is how you run it against real data.
Proceed directly with the task — run the ax command you need. Do NOT check versions, env vars, or profiles upfront.
If an ax command fails, troubleshoot based on the error:
command not found or version error → see references/ax-setup.md401 Unauthorized / missing API key → run ax profiles show to inspect the current profile. If the profile is missing or the API key is wrong, follow references/ax-profiles.md to create/update it. If the user doesn't have their key, direct them to https://app.arize.com/admin > API Keysax spaces list to pick by name, or ask the userax ai-integrations list --space SPACE to check for platform-managed credentials. If none exist, use the arize-ai-provider-integration skill — never ask the user to paste a provider key into chat..env files or search the filesystem for credentials. Use ax profiles for Arize credentials and ax ai-integrations for LLM provider keys. Never ask the user to paste secrets into chat. For missing credentials, see references/ax-profiles.md.ax ai-integrations list, (4) contact support at https://arize.com/supportAn evaluator is an LLM-as-judge definition. It contains:
| Field | Description |
|---|---|
| Template | The judge prompt. Uses {{variable}} (double-brace) placeholders (e.g. {{input}}, {{output}}, {{context}}) that get filled in at run time via a task's column mappings. |
| Classification choices | The set of allowed output labels (e.g. factual / hallucinated). Binary is the default and most common. Each choice can optionally carry a numeric score. |
| AI Integration | Stored LLM provider credentials (OpenAI, Anthropic, Bedrock, etc.) the evaluator uses to call the judge model. |
| Model | The specific judge model (e.g. gpt-4o, claude-sonnet-4-5). |
| Invocation params | Optional JSON of model settings like {"temperature": 0}. Low temperature is recommended for reproducibility. |
| Optimization direction | Whether higher scores are better (MAXIMIZE) or worse (MINIMIZE). Sets how the UI renders trends. |
| Data granularity | Whether the evaluator runs at the span, trace, or session level. Most evaluators run at the span level. |
Evaluators are versioned — every prompt or model change creates a new immutable version. The most recent version is active.
Code evaluators are the deterministic alternative — no AI integration or model, just Python. They run as a class subclassing CodeEvaluator, not a bare function, and have their own strict import-path and evaluate()-signature contract. Getting either wrong makes a run cancel silently at 0/0/0. See "Custom Python code evaluators" in references/cli-reference.md before writing one.
A task is how you run one or more evaluators against real data. Tasks are attached to a project (live traces/spans) or a dataset (experiment runs). A task contains:
| Field | Description |
|---|---|
| Evaluators | List of evaluators to run. You can run multiple in one task. |
| Column mappings | Maps each evaluator's template variables to actual field paths on spans or experiment runs (e.g. "input" → "attributes.input.value"). This is what makes evaluators portable across projects and experiments. |
| Query filter | SQL-style expression to select which spans/runs to evaluate (e.g. "span_kind = 'LLM'"). Optional but important for precision. |
| Continuous | For project tasks: whether to automatically score new spans as they arrive. |
| Sampling rate | For continuous project tasks: fraction of new spans to evaluate (0–1). |
Set --data-granularity when creating the evaluator (ax evaluators create-template-evaluator / create-code-evaluator), not on the task — it controls what unit of data that evaluator scores whenever it runs against a project task (not dataset/experiment tasks — those evaluate experiment runs directly). It defaults to span.
| Level | What it evaluates | Use for | Result column prefix |
|---|---|---|---|
span (default) | Individual spans | Q&A correctness, hallucination, relevance | eval.{name}.label / .score / .explanation |
trace | All spans in a trace, grouped by context.trace_id | Agent trajectory, task correctness — anything that needs the full call chain | trace_eval.{name}.label / .score / .explanation |
session | All traces in a session, grouped by attributes.session.id and ordered by start time | Multi-turn coherence, overall tone, conversation quality | session_eval.{name}.label / .score / .explanation |
For trace granularity, spans sharing the same context.trace_id are grouped together. Column values used by the evaluator template are comma-joined into a single string (each value truncated to 100K characters) before being passed to the judge model.
For session granularity, the same trace-level grouping happens first, then traces are ordered by start_time and grouped by attributes.session.id. Session-level values are capped at 100K characters total.
{{conversation}} template variableAt session granularity, {{conversation}} is a special template variable that renders as a JSON array of {input, output} turns across all traces in the session, built from attributes.input.value / attributes.llm.input_messages (input side) and attributes.output.value / attributes.llm.output_messages (output side).
At span or trace granularity, {{conversation}} is treated as a regular template variable and resolved via column mappings like any other.
Note: For
{{conversation}}to work, spans must carryattributes.session.id. See the arize-instrumentation skill for how to emitsession.idfrom application code, including theforce_flush()pattern required for Jupyter notebooks and short-lived scripts.
A task can contain evaluators at different granularities. At runtime the system uses the highest granularity (session > trace > span) for data fetching and automatically splits into one child run per evaluator. Per-evaluator query_filter in the task's evaluators JSON further narrows which spans are included (e.g., only tool-call spans within a session).
Full command reference for AI integrations, evaluators (template and code), and tasks — every flag, with examples — is in references/cli-reference.md. The workflows below include the commands you need inline.
Use this when the user says something like "create an evaluator for my Playground Traces project".
ax spans export accepts a project name directly — no ID lookup needed. If you don't know the project name, list available projects:
Find the entry whose "name" matches (case-insensitive) and use that name as PROJECT in subsequent commands. If you later hit a validation error with a name, fall back to using the project's "id" (a base64 string) instead.
If the user specified the evaluator type (hallucination, correctness, relevance, etc.) → skip to Step 3.
If not, sample recent spans to base the evaluator on actual data:
Inspect attributes.input, attributes.output, span kinds, and any existing annotations. Identify failure modes (e.g. hallucinated facts, off-topic answers, missing context) and propose 1–3 concrete evaluator ideas. Let the user pick.
Each suggestion must include: the evaluator name (bold), a one-sentence description of what it judges, and the binary label pair in parentheses. Format each like:
label_a / label_b)Example:
correct / incorrect)factual / hallucinated)If a suitable integration exists, note its ID. If not, create one using the arize-ai-provider-integration skill. Ask the user which provider/model they want for the judge.
Use the template design best practices below. Keep the evaluator name and variables generic — the task (Step 6) handles project-specific wiring via column_mappings.
Recommended approach: Always start with a small backfill (~100 historical spans) to validate the evaluator before turning on continuous monitoring. This lets you catch column mapping errors, wrong span kinds, and template issues on known data before scoring all future production spans. Only enable continuous after a backfill confirms correct scoring.
Before creating the task, ask:
"Would you like to: (a) Run a backfill on historical spans (one-time)? (b) Set up continuous evaluation on new spans going forward? (c) Both — backfill first to validate, then keep scoring new spans automatically? (recommended)"
Do not guess paths. Pull a sample and inspect what fields are actually present:
For each template variable ({{input}}, {{output}}, {{context}}), find the matching JSON path. Common starting points — always verify on your actual data before using:
| Template var | LLM span | CHAIN span |
|---|---|---|
input | attributes.input.value | attributes.input.value |
output | attributes.llm.output_messages.0.message.content | attributes.output.value |
context | attributes.retrieval.documents.contents | — |
tool_output | attributes.input.value (fallback) | attributes.output.value |
Validate span kind alignment: If the evaluator prompt assumes LLM final text but the task targets CHAIN spans (or vice versa), runs can cancel or score the wrong text. Make sure the query_filter on the task matches the span kind you mapped.
query_filter only works on indexed attributes: The query_filter in the evaluators JSON is evaluated against the eval index, not the raw span store. Attributes under attributes.metadata.* or custom keys may not be indexed and will silently match nothing. Use well-known indexed attributes like span_kind or attributes.llm.model_name for filtering. If a filter returns 0 spans despite data existing, try removing the filter as a diagnostic step.
Full example --evaluators JSON:
Include a mapping for every variable the template references. Omitting one causes runs to produce no valid scores.
Backfill only (a):
Continuous only (b):
Both (c): Use --is-continuous on create, then also trigger a backfill run in Step 8.
Eval index lag: The eval index is built asynchronously from the primary trace store and can lag 1–2 hours. For your first test run, use a time window ending at least 2 hours in the past. If you set
--data-end-timeto "now" on spans ingested in the last hour, the run will complete successfully but score 0 spans.
First find what time range has data:
Use the start_time / end_time fields from real spans to set the window. For the first validation run, cap --max-spans at ~100 to get quick feedback:
Review scores and explanations before widening to the full backfill or enabling continuous.
Use this when the user says something like "create an evaluator for my experiment" or "evaluate my dataset runs".
If the user says "dataset" but doesn't have an experiment: A task must target an experiment (not a bare dataset). Ask:
"Evaluation tasks run against experiment runs, not datasets directly. Would you like help creating an experiment on that dataset first?"
If yes, use the arize-experiment skill to create one, then return here.
Note the dataset name and the experiment name(s) to score. These accept names or IDs in subsequent commands — names are preferred.
If the user specified the evaluator type → skip to Step 3.
If not, inspect a recent experiment run to base the evaluator on actual data:
Look at the output, input, evaluations, and metadata fields. Identify gaps (metrics the user cares about but doesn't have yet) and propose 1–3 evaluator ideas. Each suggestion must include: the evaluator name (bold), a one-sentence description, and the binary label pair in parentheses — same format as Workflow A, Step 2.
Same as Workflow A, Step 3.
Same as Workflow A, Step 4. Keep variables generic.
Run data shape differs from span data. Inspect:
Common mapping for experiment runs:
output → "output" (top-level field on each run)input → check if it's on the run or embedded in the linked dataset examplesIf input is not on the run JSON, export dataset examples to find the path:
--experiment-ids takes the base64 ID from ax experiments list --space SPACE -o json.
Use {{input}}, {{output}}, and {{context}} — not names tied to a specific project or span attribute (e.g. do not use {{attributes_input_value}}). The evaluator itself stays abstract; the task's column_mappings is where you wire it to the actual fields in a specific project or experiment. This lets the same evaluator run across multiple projects and experiments without modification.
Use exactly two clear string labels (e.g. hallucinated / factual, correct / incorrect, pass / fail). Binary labels are:
If the user insists on more than two choices, that's fine — but recommend binary first and explain the tradeoff (more labels → more ambiguity → lower inter-rater reliability).
The template must tell the judge model to respond with only the label string — nothing else. The label strings in the prompt must exactly match the labels in --classification-choices (same spelling, same casing).
Good:
Bad (too open-ended):
Pass --invocation-params '{"temperature": 0}' for reproducible scoring. Higher temperatures introduce noise into evaluation results.
--include-explanations for debuggingDuring initial setup, always include explanations so you can verify the judge is reasoning correctly before trusting the labels at scale.
Single quotes prevent the shell from interpolating {{variable}} placeholders. Double quotes will cause issues:
--classification-choices to match your template labelsThe labels in --classification-choices must exactly match the labels referenced in --template (same spelling, same casing). Omitting --classification-choices causes task runs to fail with "missing rails and classification choices."
| Problem | Solution |
|---|---|
ax: command not found | See references/ax-setup.md |
401 Unauthorized | API key may not have access to this space. Verify at https://app.arize.com/admin > API Keys |
Evaluator not found | ax evaluators list --space SPACE |
Integration not found | ax ai-integrations list --space SPACE |
Task not found | ax tasks list --space SPACE |
project and dataset-id are mutually exclusive | Use only one when creating a task |
experiment-ids required for dataset tasks | Add --experiment-ids to create and trigger-run |
sampling-rate only valid for project tasks | Remove --sampling-rate from dataset tasks |
Validation error on ax spans export | Project name usually works; if you still get a validation error, look up the base64 project ID via ax projects list --space SPACE -o json and use the id field instead |
| Template validation errors | Use single-quoted --template '...' in bash; double braces {{var}}, not single {var} |
Run stuck in pending | ax tasks get-run RUN_ID; then ax tasks cancel-run RUN_ID |
Run cancelled ~1s | Integration credentials invalid — check AI integration |
Run cancelled ~3min | Found spans but LLM call failed — wrong model name or bad key |
Run completed, 0 spans | Widen time window; eval index may not cover older data |
| No scores in UI | Fix column_mappings to match real paths on your spans/runs |
| Scores look wrong | Add --include-explanations and inspect judge reasoning on a few samples |
| Evaluator cancels on wrong span kind | Match query_filter and column_mappings to LLM vs CHAIN spans |
Time format error on trigger-run | Use 2026-03-21T09:00:00 — no trailing Z |
| Run failed: "missing rails and classification choices" | Add --classification-choices '{"label_a": 1, "label_b": 0}' to ax evaluators create-template-evaluator — labels must match the template |
Run completed, all spans skipped | Query filter matched spans but column mappings are wrong or template variables don't resolve — export a sample span and verify paths |
query_filter set but 0 spans scored | The filter attribute may not be indexed in the eval index. attributes.metadata.* and custom attributes are often not indexed. Use span_kind or attributes.llm.model_name instead, or remove the filter to confirm spans exist in the window. |
Custom code evaluator run cancels ~3s with 0/0/0 (successes/errors/skipped) | Wrong import path or evaluate() signature — see the "CRITICAL" callout under Custom Python code evaluators in references/cli-reference.md. Must import from arize.experimental.datasets.experiments.evaluators.base (not arize.experiments) and declare named evaluate() params, not just **kwargs. |
When a task run reports status cancelled, work through the ordered checklist (credentials → model name → column-mapping/path checks → time window → span kind → variable resolution) in references/troubleshooting.md.
See references/ax-profiles.md § Save Credentials for Future Use.**