All resources

Everything we’ve published — page 9.

Blog

How to build LLM-as-a-Judge evaluators that hold up in production

Learn how to design, calibrate, and run LLM-as-a-judge evaluators with fixed labels, human agreement checks, trace context, and…

Read the post
Blog

What we learned testing 7 models under the same agent harness

Model swaps look like configuration changes, but they behave more like product migrations. A new model may be…

Read the post
Guide

Braintrust alternatives for AI observability & agent evaluations

Compare the top Braintrust alternatives for AI observability and evaluation, focusing on tools that support tracing, debugging, and…

Read the guide
Blog

Building a self-improving agent on a context graph of human disagreement

You can build a measurably better agent from data you already have, without retraining a thing. The data…

Read the post
Post

Coding agent tracing and evaluation: An open source tool to improve AI coding workflows

Announcing coding harness tracing for observing, evaluating, and improving coding agent workflows across Claude Code, Cursor, Codex, GitHub…

Read the post
Blog

How we use Alyx to build Alyx: How to build an AI agent feedback loop

How Arize uses Alyx to debug Alyx: searching dense traces, aggregating failures, triaging dogfooding issues, and closing the…

Read the post
Guide

AI agent analytics platforms: A buyer’s guide

What to measure, how to evaluate, and what to require when buying analytics for production AI agents—spans, traces,…

Read the guide
Blog

Models got an order of magnitude better at following instructions in one year

A year ago, frontier models started losing track of instructions somewhere around 200–300 simultaneous constraints. With 2026 models,…

Read the post
Post

From observability to context: What’s next for Arize Phoenix

As agents start changing software, they need a way to verify their work that includes traces, evals, feedback,…

Read the post
Blog

Agent harnesses have an expiration date

A benchmark-driven look at why agent harnesses need adaptive finish logic as model behavior changes across Claude, GPT-4o,…

Read the post
Blog

AI agent evaluation: How to test, debug, and improve agents in production

Lessons from building and shipping Alyx, our AI agent

Read the post
Post

Swarm management in agent harnesses: owning long-running agents

As we have built our own harness management tools internally at Arize, and watched external systems like Devin…

Read the post

Don’t ship vibes.

Arize gives AI teams observability and evals to understand and improve agent performance.