-
AI EvaluationHow to build LLM-as-a-Judge evaluators that hold up in production
Learn how to design, calibrate, and run LLM-as-a-judge evaluators with fixed labels, human agreement checks, trace context, and Phoenix Evals. Aaron Winston May 21, 2026 22 min read -
AI EvaluationWhat we learned testing 7 models under the same agent harness
Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on… Nancy Chauhan May 20, 2026 10 min read -
AI EvaluationCoding agent tracing and evaluation: An open source tool to improve AI coding workflows
Announcing coding harness tracing for observing, evaluating, and improving coding agent workflows across Claude Code, Cursor, Codex, GitHub Copilot, and Gemini CLI. Duncan McKinnon Chris Cooning Fuad Ali May 18, 2026 5 min read -
AI EvaluationHow we use Alyx to build Alyx: How to build an AI agent feedback loop
How Arize uses Alyx to debug Alyx: searching dense traces, aggregating failures, triaging dogfooding issues, and closing the AI engineering feedback loop. Chris Cooning Sally-Ann DeLucia Priyan Jindal Jack Zhou May 13, 2026 10 min read -
AI EvaluationModels got an order of magnitude better at following instructions in one year
A year ago, frontier models started losing track of instructions somewhere around 200–300 simultaneous constraints. With 2026 models, that ceiling is closer to 2,000 — an order-of-magnitude jump.… Laurie Voss May 12, 2026 11 min read -
AI EvaluationFrom observability to context: What’s next for Arize Phoenix
As agents start changing software, they need a way to verify their work that includes traces, evals, feedback, and APIs. This is where Phoenix goes next — not… Mikyo King Elizabeth Hutton May 11, 2026 10 min read -
AI EvaluationAI agent evaluation: How to test, debug, and improve agents in production
Lessons from building and shipping Alyx, our AI agent Sally-Ann DeLucia Chris Cooning Priyan Jindal Jack Zhou May 5, 2026 9 min read -
AI EvaluationWhat is an evaluation harness? Definition & guide
An evaluation harness is the standardized infrastructure that decides what gets evaluated, runs the evaluation, and acts on the result. Chris Cooning Hakan Tekgul Cam Young May 4, 2026 14 min read -
AI EvaluationMCP vs. CLI Skills for agents: what our eval found (and which you should use)
Twitter said pick a side. The eval said the question was wrong. Six months ago, MCP (model context protocol) was the hot new thing: tool usage with a… Laurie Voss May 1, 2026 10 min read -
AI EvaluationBeyond models: How context and evals make agents work in production
Building an AI agent has never been easier. But getting one into production that’s reliable is still hard. Most teams can ship a working demo in a day.… Patrick Kelly April 23, 2026 9 min read -
AI EvaluationHow to add an evaluation harness to your Gemini CLI coding agent
Coding agents can update prompts, wire in tools, and change application logic across your codebase in a single run. The hard part isn’t getting the agent to make… Richard Young April 22, 2026 7 min read -
AI EvaluationCode is free, technical debt isn’t: Notes from AI Engineer Europe
Keynotes at Europe’s first flagship AI Engineer Conference shared one theme: code generation has accelerated past our ability to verify it, and the industry is quietly reorganizing around… RL Nabors April 20, 2026 5 min read -
AI EvaluationBuilding smarter AI agents: architecture, evals, and lessons from the field
Shipping an AI agent is easy. Understanding whether it actually works in production is not. That was the common thread across two AI Builders events in San Francisco… Jim Bennett April 14, 2026 10 min read -
AI EvaluationFrom First Eval to Autonomous AI Ops: A Maturity Model for AI Evaluation
Every team runs evals. Almost none have an evaluation practice. The difference is the gap between a one-off notebook and a system that continuously assesses, alerts, and acts… Cam Young April 3, 2026 6 min read -
AI EvaluationHow We Used Evals (and an AI Agent) to Iteratively Improve an AI Newsletter Generator
We love building little AI-powered tools that accelerate our workflows. One we built recently is a tool that takes our recent tweets and uses Claude to create a… Laurie Voss March 10, 2026 10 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.