Evaluate production traces with Jev-as-a-Judge directly in Arize AX

Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow.

Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow.

Jev is TypeSafe AI’s System One decision model, built for structured decisions rather than text generation. With the new integration, teams can evaluate traces with boolean, choice, and score-based questions and get structured labels and confidence scores back directly in Arize AX.

A faster option for bounded evaluation tasks

Not every evaluation needs a general-purpose LLM to reason through a response and generate an explanation.

Many production evals ask relatively bounded questions:

  • Did the agent resolve the user’s request?
  • Did it choose the right tool?
  • Does this response satisfy a defined policy?
  • Which category does this interaction belong to?

Jev is designed for this kind of classification, scoring, and routing. You provide a shared state from the trace and one or more typed questions, and Jev returns structured results in a single call.

That gives AI teams another tool for matching the evaluator to the job. You can start by testing Jev on evaluations with explicit criteria and a fixed set of possible outcomes. Compare its judgment with human labels on representative examples before deciding whether its accuracy, cost, and latency meet your requirements. An LLM judge may be useful when you need a written explanation, while an agent judge can gather additional context or investigate before reaching a verdict.

Lower evaluation costs can make it practical to assess more production traffic, giving teams more examples of where agents fail and what to improve next.

What we learned testing Jev

We tested Jev on hallucination detection to compare its accuracy, cost, and latency with LLM judges.

In one Arize benchmark for hallucination detection, a threshold-tuned Jev matched Claude Opus 5 at 87% accuracy, while running at roughly 1/300 of the cost and 23x the speed in that test. The important caveat was threshold tuning: Jev’s default cutoff performed materially worse, reinforcing the importance of calibrating evaluators against human-labeled data before deploying them broadly.

We’ve also explored Jev for real-time guardrails and previously showed how to connect it to AX using a remote evaluator. Native Jev-as-a-Judge support removes that additional service layer and brings the workflow directly into AX.

Go deeper on Jev

We’ve been exploring where decision models like Jev fit into the AI evaluation stack. Catch up on the latest Arize research and tutorials:

Get started

To use Jev-as-a-Judge in Arize AX:

  1. Add your TypeSafe AI API key under Settings → AI Providers → TypeSafe AI.
  2. Create a new Jev-as-a-Judge evaluator.
  3. Define the trace data Jev should evaluate and the questions it should answer.
  4. Test the evaluator against human-labeled examples from your application.
  5. Attach it to an online evaluation task to evaluate incoming production traces.

Follow the Jev-as-a-Judge setup guide in our docs.

You can also learn how to set up the TypeSafe AI integration in our docs.

Get the latest on AI & Observability

Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:

Don’t ship vibes.

Arize gives AI teams observability and evals to understand and improve agent performance.