The Evaluator

Your go-to blog for insights on AI observability and evaluation.

Showing 1–10 of 458 posts (page 1 of 46)

You chose the best model. Why is your agent still failing?
AI Engineering Evaluations Observability & tracing

You chose the best model. Why is your agent still failing?

Public benchmarks can show how a model performs in general. Production reliability depends on the context and harness around it, which only your team can evaluate against its own data, workflows, and users.

Arize AX adds native support for OpenTelemetry GenAI semantic conventions
Agent Observability Product Releases

Arize AX adds native support for OpenTelemetry GenAI semantic conventions

Arize AX now normalizes OpenTelemetry GenAI semantic conventions into first-class AI traces, unlocking evaluations, token and cost visibility, and easier debugging.

Demystifying the EU AI Act for AI product and engineering teams
AI Engineering Security & Governance

Demystifying the EU AI Act for AI product and engineering teams

An engineering guide to turning EU AI Act principles into traces, evaluations, annotations, and release evidence product and engineering teams can actually demonstrate.

Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:

How cheap models changed multi-agent economics
Agent Engineering Agent Evaluation Agents

How cheap models changed multi-agent economics

Orchestrator-executor just became the smart default for production agents: an expensive model plans, cheap models execute, and cost per completed task decides the roster.

AI agent observability: Why production systems need a reasoning layer
Agent Engineering Agent Observability AI Engineering

AI agent observability: Why production systems need a reasoning layer

Traditional APM can collect every span and still leave developers guessing about intent, causality, and drift. As agents multiply, the observability stack must learn to interpret the systems it watches.

How to debug production AI agents with Signal in Arize AX
AI Engineering Product tutorials

How to debug production AI agents with Signal in Arize AX

Learn how Arize Signal turns production traces into ranked issues, proposed fixes, regression datasets, and reviewable pull requests for AI agents.

Hamel Husain explains why AI evals fail before the evaluation begins
Agent Engineering Agent Evaluation Agents

Hamel Husain explains why AI evals fail before the evaluation begins

Hamel Husain explains why ambiguous inputs, generic metrics, and disconnected review workflows can make AI evaluations misleading, and how developers can build a better process around real production data.

From Signal to PR: What if your agents got better every time they failed?
Agent Engineering Agent Observability AI Engineering

From Signal to PR: What if your agents got better every time they failed?

Signal, a managed agent built into Arize AX, continuously reviews production traces, surfaces ranked issues with evidence and proposed fixes, and — with Managed Agents — can carry investigations into your codebase and open pull requests.

How to improve agent skills with tracing and evals
Agent Engineering Agent Evaluation Agent Observability

How to improve agent skills with tracing and evals

A skill cut agent costs by 44% and latency by 56%, but also reduced answer completeness. Here’s how tracing, evals, and a long-running agent exposed and corrected the regression.

Tips from Anthropic on building agent evals you can trust
Agent Engineering Agent Evaluation Agent Observability

Tips from Anthropic on building agent evals you can trust

Learn how to build trustworthy AI agent evals using regression tests, capability evals, production traces, LLM judges, and reproducible environments.