Anthropic says it fixed Claude’s writing. I ran the evals to check.
Opus 5.5 dropped em dashes from 12.9 per 1,000 words to two in 57,000 words, and cut the rest of the Claudisms in half. Better, but not fixed.
12 min read
Evaluate production traces with Jev-as-a-Judge directly in Arize AX
Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow.
3 min read
What Jev’s probabilities reveal that repeated LLM judgments miss
I tested Jev and five LLM judges across ten Arize Phoenix evaluators to measure how often their answers changed alongside accuracy, cost, and latency.
11 min readBenchmarks, agent evaluations, and experiments.
View all research
Are agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment.
September 29, 2026 9 min read
Anthropic says it fixed Claude’s writing. I ran the evals to check.
Opus 5.5 dropped em dashes from 12.9 per 1,000 words to two in 57,000 words, and cut the rest of the Claudisms in half. Better, but not fixed.
September 29, 2026 12 min read
What Jev’s probabilities reveal that repeated LLM judgments miss
I tested Jev and five LLM judges across ten Arize Phoenix evaluators to measure how often their answers changed alongside accuracy, cost, and latency.
September 28, 2026 11 min readDemos, workshops, and conference talks.
See all on YouTube
The Model Knew What to Say. The Harness Failed. | Laurie Voss
Laurie Voss (Head of Developer Relations at Arize AI) breaks down why the agent harness, specifically code-level guardrails, is the most critical boundary layer in modern software engineering.
InsightsGet the latest on AI & Observability
Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:
Real teams, shipping AI.
See allHow Booking.com scales AI observability with Arize
How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and…
Read moreHow LG Uplus is building better AI customer service agents with evaluation-driven development
How LG Uplus uses Arize AX to build evaluation-driven AI contact center agents — combining production traces, user feedback, and domain expertise to continuously improve customer service for…
Read moreHow Tripadvisor is building the AI product development lifecycle for agentic travel
Tripadvisor VP of Data and AI Rahul Todkar on building a production AI lifecycle for traditional ML and agentic travel products with unified observability, traces, evals, and governance…
Read moreCatch up on everything you missed.
See all-
Agent EngineeringHow we built UI Code Mode into Arize Phoenix
We replaced the 58 tools PXI used to drive the Phoenix UI with two tools and a JavaScript sandbox that runs in your browser tab. Here's why, how… Anthony Powell September 25, 2026 25 min read -
Agent EvaluationHow to build a remote evaluator in Arize AX with Jev
Evaluate agent responses with Jev and Arize AX. Build a FastAPI remote evaluator, return labels and scores, and choose a threshold for your data. Sally-Ann DeLucia Jim Bennett September 24, 2026 12 min read -
Agent EvaluationJev vs. LLM-as-a-Judge: Accuracy and cost benchmarks
We benchmarked Jev against Claude Opus 5 and GPT-5.6 Terra on accuracy, cost, and latency. Learn how threshold tuning changes hallucination detection. Laurie Voss September 23, 2026 11 min read -
Agent EvaluationReal-time LLM guardrails with Jev: comparing latency and cost
Compare Jev and GPT-5.4 nano for real-time LLM guardrails, with demo results on latency, cost, and checks on agent inputs, replies, and tool calls. Jim Bennett September 23, 2026 13 min read -
Agent EngineeringWhat changes when AI agents use your software
Daytona cofounder Ivan Burazin wants agents that can finish the job within the authority they have been given. His interview offers a starting point for examining how agents… Aaron Winston September 23, 2026 9 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.