Nobody agrees yet on what to call Jev, but almost everyone agrees it’s one of the most exciting model releases this month. It reads natural language like an LLM, but instead of generating text it returns a probability for each label. That matters because most LLM judges are already classifiers in disguise: we force their answers into a fixed set of labels, then validate them with the same metrics we’d use for any classifier.
My colleagues have already shown that Jev can match LLM judges on accuracy at a fraction of their cost and latency. I wanted to test something else. I’ve argued before that an LLM judge’s nondeterminism is a signal, not just a problem. When developing a new eval, I recommend running the judge ten or more times at high temperature. It wavers most on hard examples and ambiguous criteria, which shows you where to focus human review.
But that’s too expensive at production scale. Could a judge that returns a probability give you the same signal from one call?
I ran five LLMs and two Jev configurations over our existing benchmarks for 10 Phoenix evaluators, ten times each, for 36,190 judgments in all. Jev’s probability turned out to be a strong predictor of its own correctness: its low-probability answers were far more likely to be wrong than its high-probability ones. Its uncertainty also correlated moderately with variation in the LLM responses. And finally, Jev performed the same with Phoenix’s existing prompts as with its native API.
Build better agents with Arize
Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.
Prefer open source? Try Arize Phoenix for self-hosted, open source agent observability.
How we tested LLM judge consistency in Arize Phoenix
We already maintain human-labeled benchmarks for every evaluator built into Phoenix: correctness, PII detection, completeness, tool invocation, tool response handling, retrieval relevance, user friction, hallucination, refusal and conciseness.
Altogether there are 517 labeled examples, 23 to 136 per evaluator. I swept all ten benchmarks across seven judges: three small LLMs, two frontier LLMs and two variants of Jev.
Phoenix evaluator benchmarks and label distributions
| Evaluator | Examples | Label counts | Smaller class |
|---|---|---|---|
| Correctness | 79 | 39 correct, 40 incorrect | 49% |
| PII detection | 136 | 136 PII detected | none |
| Completeness | 45 | 17 complete, 28 incomplete | 38% |
| Tool invocation | 23 | 7 correct, 16 incorrect | 30% |
| Tool response handling | 50 | 24 correct, 26 incorrect | 48% |
| Retrieval relevance | 33 | 16 relevant, 17 irrelevant | 48% |
| User friction | 40 | 16 friction, 24 no friction | 40% |
| Hallucination | 37 | 18 grounded, 19 hallucinated | 49% |
| Refusal | 40 | 17 answered, 23 refused | 43% |
| Conciseness | 34 | 9 concise, 25 verbose | 26% |
| Total | 517 |
The PII detection suite contains only examples with PII, so it measures how often a judge catches PII, not how often it raises false alarms.
Jev and LLM judges compared by model class and token pricing
| Judge | Class | $ per M input | $ per M output |
|---|---|---|---|
| Jev 1.13.0 | Classifier | 0.042 | free |
| GPT-5 nano | Small | 0.05 | 0.40 |
| Gemini 3.8 Flash | Small | 0.75 | 3.75 |
| Claude Haiku 4.5 | Small | 1.00 | 5.00 |
| GPT-5.6 Sol | Frontier | 4.00 | 20.00 |
| Claude Opus 5 | Frontier | 5.00 | 25.00 |
Every judge scored every example ten times, because LLMs often change their responses with repetition. Phoenix tracked latency, tokens and cost for every run. The whole sweep cost $169.84, and Jev’s share was only 64 cents.
I report accuracy averaged across the ten evaluators, so each one counts equally. I also calculated macro F1, which weights both labels equally, but I left it out of this report because it ranks the judges in the same order and stays within 0.2 points of accuracy on the nine evaluators where it’s defined.
Drop-in and native Jev configurations
Phoenix’s built-in evaluators use prompt templates written for LLMs, with the rubric, the data to judge and the allowed labels together in one block of text. Jev’s API splits a request into three fields instead. state holds the data to judge, instructions says what to decide, and criteria lists the allowed labels with a description of each.
So I ran Jev two ways. As a drop-in, it gets the exact prompt we’d send an LLM, all of it in state. Used natively, the same rubric is split across the three fields.
How drop-in and native Jev configurations structure the same prompt
| Field | Drop-in | Native |
|---|---|---|
state |
The full Phoenix prompt, rendered with the example, as one string | Only the example’s data, as structured fields such as input and output |
instructions |
One fixed line: “Select the label that best satisfies the complete rubric in the state.” | The Phoenix rubric, with the data block taken out |
criteria |
The label names, with no descriptions | Each label, with one sentence taken from the rubric |
Drop-in is how phoenix-evals runs Jev today, treating it like any other LLM. I wanted to know whether that leaves accuracy on the table compared with Jev’s native API.
Drop-in vs. native Jev results across 10 evaluators
Averaged across the 10 evaluators, Jev scored 98.3% accuracy both as a drop-in and native, and none of the per-task differences between the two was significant. So from here on, “Jev” means the drop-in setup, since that’s how phoenix-evals runs it.
The differences that did show up follow a pattern. Native’s biggest gain, 7.1 points, came on completeness, which asks the judge to check every request in a conversation averaging around 1,950 tokens.
Pulling the data out of that much instruction text seems to help. Its losses on retrieval relevance and tool invocation come down to one or two examples each. If your evaluator reads long inputs, native is worth a try.
How often Jev and LLM judges changed their answers
A judge that returns a different label for the same input moves your eval scores with nothing in your app changing. Here, a judge changed its answer when at least one of an example’s ten runs disagreed with the rest.
Jev changed its answer on 0.97% of examples as a drop-in and 0.19% native, the lowest of any judge. Comparing example by example, drop-in Jev was significantly steadier than three of the five LLMs, including GPT-5 nano and Haiku 4.5, and level with Gemini 3.8 Flash and Opus 5.
The worst cases were GPT-5 nano on tool invocation, where it changed its answer on 26% of examples, and Haiku 4.5 on PII detection, at 12%. Jev isn’t fully deterministic, since its probabilities move a little between calls, but they rarely move enough to change the label.
Can Jev’s probabilities flag evaluation errors?
Jev returns a probability with every label, so the first thing I checked was whether that number drops when Jev gets the answer wrong. It does, sharply.
On drop-in Jev’s correct answers, the median probability on the chosen label was 1.00. On its wrong answers, it was 0.71.
All 3,561 runs where Jev put at least 99% on its answer were correct. Below 70%, it was right only 73% of the time. As a detector for its own mistakes, the probability scores an ROC AUC of 0.95 as a drop-in and 0.99 native.
In this comparison, the LLM judges returned labels without per-label probabilities. I used agreement across ten repeated calls as a separate uncertainty signal.
Measured that way, the LLMs’ ROC AUC for spotting the examples they got wrong ranged from 0.62 to 0.84, against 0.95 for Jev. Many LLM mistakes are consistent. Of the examples an LLM got wrong on most runs, 27 to 75% never wavered at all, depending on the model, so repetition would never have caught them.
Jev was just as steady on its own mistakes, but its probability still told you it was unsure. Each judge got only 4 to 14 examples wrong, so treat these numbers as a strong direction, not a precise ranking.
Using Jev’s uncertainty to prioritize human review
The same probabilities also point at examples that are hard for everyone. I compared Jev’s uncertainty with whether any of the five LLM judges changed its answer on the same example across ten runs. I measured Jev’s uncertainty as entropy. Zero bits means all its probability is on one label, and one bit means a coin flip.
On the 306 examples where Jev was nearly certain, an LLM judge wavered on 3%. On the 24 where Jev was close to a coin flip, the LLMs wavered on 71%. The rank correlation is 0.47, and it survives a test that shuffles values within each evaluator, so it isn’t just separating easy evaluators from hard ones.
As a detector for examples that make LLM judges unstable, Jev’s entropy has an ROC AUC of 0.85. Given one such example and one stable one, it rates the unstable one as more uncertain 85% of the time.
Jev’s probabilities tell you which examples are hard, without running an LLM ten times to find out. A 0.5-bit cutoff flags 11% of examples, and the LLMs wavered on nearly two thirds of them. Those are the exact subset of examples you’d want a human to review.
Accuracy across 10 evaluators is a tie, but cost and latency aren’t
With 23 to 136 examples per evaluator, no pairwise difference in accuracy between any two judges held up after correcting for multiple comparisons, and that includes every comparison involving Jev. A test across all seven judges at once found a real difference on only two evaluators, completeness and PII detection, and on both it came mostly from Claude Haiku 4.5 trailing the rest.
Jev and LLM judges compared on accuracy, consistency, cost, and latency
| Judge | Accuracy | Changed answer | $ per 1,000 judgments | Median latency |
|---|---|---|---|---|
| Jev, drop-in | 98.3% | 0.97% | $0.06 | 133 ms |
| Jev, native | 98.3% | 0.19% | $0.06 | 133 ms |
| GPT-5 nano | 96.7% | 7.9% | $0.71 | 7.3 s |
| Claude Haiku 4.5 | 97.1% | 6.8% | $2.27 | 2.3 s |
| Gemini 3.8 Flash | 98.6% | 1.2% | $2.32 | 1.7 s |
| GPT-5.6 Sol | 98.0% | 3.3% | $8.08 | 2.4 s |
| Claude Opus 5 | 99.4% | 0.77% | $17.08 | 4.0 s |
Accuracy is averaged across the ten evaluators. Changed answer is the share of examples where at least one of ten runs returned a different label.
Directionally and statistically, Jev landed with the pack. Its 98.3% sits inside the small models’ range of 96.7 to 98.6% and the frontier models’ range of 98.0 to 99.4%. Claude Opus 5 had the highest average, but its lead over Jev wasn’t significant on any evaluator.
Cost and speed don’t need a significance test. Jev costs 12 to 40 times less than the small models and about 140 to 290 times less than the frontier ones, and its 133 ms median latency is 13 to 55 times faster than any LLM here.
Accuracy by evaluator and judge
| Evaluator | Jev, drop-in | Jev, native | GPT-5 nano | Claude Haiku 4.5 | Gemini 3.8 Flash | GPT-5.6 Sol | Claude Opus 5 |
|---|---|---|---|---|---|---|---|
| Correctness | 95.2% | 96.2% | 90.8% | 95.7% | 95.8% | 87.5% | 95.7% |
| Tool response handling | 96.6% | 96.0% | 94.4% | 95.8% | 96.8% | 95.0% | 98.0% |
| Completeness | 92.9% | 100.0% | 95.1% | 88.4% | 100.0% | 100.0% | 100.0% |
| Tool invocation | 100.0% | 95.7% | 97.0% | 100.0% | 96.1% | 98.3% | 100.0% |
| User friction | 100.0% | 100.0% | 95.0% | 97.0% | 100.0% | 100.0% | 100.0% |
| PII detection | 98.6% | 100.0% | 99.9% | 95.5% | 99.2% | 99.7% | 99.8% |
| Hallucination | 100.0% | 100.0% | 99.2% | 99.5% | 97.8% | 100.0% | 100.0% |
| Retrieval relevance | 100.0% | 94.8% | 97.6% | 99.1% | 100.0% | 100.0% | 100.0% |
| Refusal | 100.0% | 100.0% | 98.5% | 100.0% | 100.0% | 100.0% | 100.0% |
| Conciseness | 100.0% | 100.0% | 99.7% | 100.0% | 100.0% | 100.0% | 100.0% |
For binary classification evals like these, Jev is a strong candidate to test against your current small LLM judge.
Frontier models are a closer call. Opus 5’s biggest edges over Jev were on completeness, which native Jev closes, and on tool response handling, 98.0% against 96.6%, an agent task where the judge has to follow what the agent did with a tool’s output.
Neither gap is significant, but they point where I’d expect a frontier judge to earn its price: evals that need several steps of reasoning, and anything where you want the verdict explained in words, which Jev never does.
How to interpret the 0.5 cutoff in this benchmark
On a two-label Choice question, taking Jev’s top label is the same as a 0.5 cutoff on its probability. I checked whether a different cutoff would help. Leaving out PII, whose suite contains only positive examples, the best cutoff added about 1 to 3 points of balanced accuracy on five tasks, depending on the setup, and nothing on the other four.
Even that is optimistic, because the suites are too small to choose the cutoff on one half and score it on the other.
These small suites provide weaker evidence about threshold tuning. Our earlier RAGTruth benchmark used a different dataset and question format and benefited from a tuned cutoff.
Before adopting 0.5 on your workload, choose a cutoff against human labels and evaluate it on separate held-out data, even when your classes are balanced.
Benchmark limitations
The suites contain 23 to 136 examples each, so the accuracy comparisons have limited power to detect small differences. Repeated calls help measure consistency, but do not add new test examples. Every task uses two labels, and the PII suite contains only examples with PII, so it does not measure false positives.
These results do not establish performance on graded scores or open-ended evaluation.
How to test judge consistency on your own data
- Label 50 to 100 examples of your own traffic.
- Run your current judge and Jev over them as Phoenix experiments, with three to five repetitions each.
- Compare accuracy with its interval, and read the flip rate next to it.
- Keep Jev’s probabilities and send its uncertain examples to a person.
I’d consider switching once Jev meets my accuracy and consistency requirements on representative examples and delivers useful cost or latency savings.
You can also try this today. Jev is available as a model option for classification evaluators in phoenix-evals, where it runs as a drop-in, so the evaluator APIs stay the same.