Building production-ready AI agents with Atlas Agent Engine and Arize AX
Agents built and deployed on MongoDB’s Atlas Agent Engine can export OpenTelemetry traces to Arize AX, where teams can inspect each run, evaluate both the outcome and execution…
6 min read
What everyone was talking about at WeAreDevelopers World Congress North America
After 237 talks across eight stages, code review emerged as the bottleneck for AI-generated code. Here are the production eval, sandbox, context, and agent experience themes that kept…
16 min read
Are agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment.
9 min readBenchmarks, agent evaluations, and experiments.
View all research
Are agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment.
September 29, 2026 9 min read
Anthropic says it fixed Claude’s writing. I ran the evals to check.
Opus 5.5 dropped em dashes from 12.9 per 1,000 words to two in 57,000 words, and cut the rest of the Claudisms in half. Better, but not fixed.
September 29, 2026 12 min read
What Jev’s probabilities reveal that repeated LLM judgments miss
I tested Jev and five LLM judges across ten Arize Phoenix evaluators to measure how often their answers changed alongside accuracy, cost, and latency.
September 28, 2026 11 min readDemos, workshops, and conference talks.
See all on YouTube
The Model Knew What to Say. The Harness Failed. | Laurie Voss
Laurie Voss (Head of Developer Relations at Arize AI) breaks down why the agent harness, specifically code-level guardrails, is the most critical boundary layer in modern software engineering.
InsightsGet the latest on AI & Observability
Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:
Real teams, shipping AI.
See allHow Booking.com scales AI observability with Arize
How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and…
Read moreHow LG Uplus is building better AI customer service agents with evaluation-driven development
How LG Uplus uses Arize AX to build evaluation-driven AI contact center agents — combining production traces, user feedback, and domain expertise to continuously improve customer service for…
Read moreHow Tripadvisor is building the AI product development lifecycle for agentic travel
Tripadvisor VP of Data and AI Rahul Todkar on building a production AI lifecycle for traditional ML and agentic travel products with unified observability, traces, evals, and governance…
Read moreCatch up on everything you missed.
See all-
Agent EvaluationEvaluate production traces with Jev-as-a-Judge directly in Arize AX
Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow. Chris Cooning September 28, 2026 3 min read -
Agent EngineeringHow we built UI Code Mode into Arize Phoenix
We replaced the 58 tools PXI used to drive the Phoenix UI with two tools and a JavaScript sandbox that runs in your browser tab. Here's why, how… Anthony Powell September 25, 2026 25 min read -
Agent EvaluationHow to build a remote evaluator in Arize AX with Jev
Evaluate agent responses with Jev and Arize AX. Build a FastAPI remote evaluator, return labels and scores, and choose a threshold for your data. Sally-Ann DeLucia Jim Bennett September 24, 2026 12 min read -
Agent EvaluationJev vs. LLM-as-a-Judge: Accuracy and cost benchmarks
We benchmarked Jev against Claude Opus 5 and GPT-5.6 Terra on accuracy, cost, and latency. Learn how threshold tuning changes hallucination detection. Laurie Voss September 23, 2026 11 min read -
Agent EvaluationReal-time LLM guardrails with Jev: comparing latency and cost
Compare Jev and GPT-5.4 nano for real-time LLM guardrails, with demo results on latency, cost, and checks on agent inputs, replies, and tool calls. Jim Bennett September 23, 2026 13 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.