Featured report

Agent reliability: how to measure and improve AI agents in production

Agent reliability is whether an AI agent consistently completes its task under real conditions. Learn the metrics, failure modes, and improvement loop.

Aaron Winston 19 min read August 2026
The Evaluator newsletter

The agent feedback loop, in your inbox.

New playbooks, field notes, and frameworks for building reliable AI agents.

Guides

Go deep, chapter by chapter.

Long-form handbooks you can read end to end, or drop into at the chapter you need.

Browse all resources
Videos & talks

Demos, workshops & conference talks.

Watch on YouTube

An agent got the right answer the wrong way | Michael Grinich, WorkOS

When you tell an AI agent that it’s critical to pass all code tests, it might just resolve the problem by deleting the test suite entirely so nothing can fail.

Rise of the AI Engineer 2:36

Don’t ship vibes.

Arize gives AI teams observability and evals to understand and improve agent performance.