Agent-as-a-Judge is available in closed Enterprise beta. Contact your Arize account team for access.

An agent evaluator in Arize AX
Why use it
Agent-as-a-Judge is for subjective or complex quality checks where you want an agent to interpret production data—not just fill a template. Examples:
- Relevance or helpfulness when the right answer depends on full trace context
- Agent trajectory quality (tool choice, recovery, multi-step reasoning)
- Custom rubrics that are easier to describe in prose than to wire into
{variable}mappings
How it works
Configure the evaluator in the Evaluator Hub—Evaluators → Create → Agent-as-a-Judge. Select a harness (Claude Code is supported today; Cursor, Codex, Hermes, and OpenClaw are coming soon), pick an Anthropic model (or Auto), then write scoring instructions in plain language. Optional placeholders like{attributes.output.value} are filled from span data. Optionally define fixed labels or let the harness decide each run.
Attach to an online eval task on an LLM project—date range, query filter, and sampling rate—same flow as Run online evals on traces.
On each run, the platform starts the selected harness. The harness reads exported spans for the task window, scores them from your instructions, and publishes eval.<name>.* columns on the spans.
View results on traces, in dashboards, and in task run history. See View eval results.
The harness gets read access to traces on the bound project automatically. Add the optional Arize skill only if you need broader API access in the harness.
Create an Agent-as-a-Judge evaluator
1
Open Evaluator Hub
Go to Evaluators in the space sidebar, then Create and choose Agent-as-a-Judge.
2
Select harness
Choose a harness. Claude Code is supported today; Cursor, Codex, Hermes, and OpenClaw are coming soon.
3
Select model
Pick an Anthropic AI integration and model, or Auto.
4
Write scoring instructions
Describe what good and bad look like—for example, whether the assistant’s response is relevant to the user input given the full trace.The agent reads traces at run time; you do not map template variables to columns upfront.
5
Configure labels (optional)
Leave Let agent decide labels on for open-ended rubrics, or turn it off to define fixed labels and scores (for example
relevant / irrelevant).6
Save to Evaluator Hub
The evaluator is versioned like LLM and code evaluators—reuse it across tasks.