How Uber evaluates AI agents at production scale
A background comment about pizza exposed a failure that Uber’s offline evaluations had missed. The incident helped reveal what production AI agent evaluation actually requires: automatic tracing, living datasets, shared ownership, and a direct connection to…
Arize and Dynatrace: Making the World’s AI Work
Today we are announcing the signing of a definitive agreement for the acquisition of Arize by Dynatrace to accelerate our mission to make the world's AI work.
5 min read
AI agent guardrails vs. evals: How to build more reliable agent systems
Guardrails constrain what an agent can do in code; evals judge whether it performed well. Learn how both layers—and the harness around them—make long-running AI agents reliable.
9 min read
Evaluation-driven development: How to move AI agents from pilot to production
Learn how evaluation-driven development, agent harnesses, AI observability, guardrails, and cost-per-outcome metrics move AI agents from pilot to production.
19 min readDemos, workshops, and conference talks.
See all on YouTubeWhy AI Agents Fail at Tasks They Already Completed | Ivan Burazin, Daytona
Ivan Burazin (CEO & Co-founder of Daytona) breaks down why the model is rarely the bottleneck; tool access, broken credential boundaries, and flawed harness design are.
InsightsGet the latest on AI & Observability
Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:
Real teams, shipping AI.
See allHow Booking.com scales AI observability with Arize
How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and…
Read moreHow LG Uplus is building better AI customer service agents with evaluation-driven development
How LG Uplus uses Arize AX to build evaluation-driven AI contact center agents — combining production traces, user feedback, and domain expertise to continuously improve customer service for…
Read moreHow Tripadvisor is building the AI product development lifecycle for agentic travel
Tripadvisor VP of Data and AI Rahul Todkar on building a production AI lifecycle for traditional ML and agentic travel products with unified observability, traces, evals, and governance…
Read moreCatch up on everything you missed.
See all-
Agent ObservabilityCrew Studio launches with native Arize AX tracing and evaluation
Through a native Arize AX integration, teams can send traces from Crew Studio to Arize from the first run without custom instrumentation—then inspect behavior, evaluate quality, and test… Richard Young Jesse Miller August 13, 2026 5 min read -
AI EngineeringYou chose the best model. Why is your agent still failing?
Public benchmarks can show how a model performs in general. Production reliability depends on the context and harness around it, which only your team can evaluate against its… Aparna Dhinakaran Prukalpa Sankar August 12, 2026 12 min read -
Agent ObservabilityArize AX adds native support for OpenTelemetry GenAI semantic conventions
Arize AX now normalizes OpenTelemetry GenAI semantic conventions into first-class AI traces, unlocking evaluations, token and cost visibility, and easier debugging. Chris Cooning Dheeraj Bandaru August 11, 2026 4 min read -
AI EngineeringDemystifying the EU AI Act for AI product and engineering teams
An engineering guide to turning EU AI Act principles into traces, evaluations, annotations, and release evidence product and engineering teams can actually demonstrate. Jitendra Yadav August 10, 2026 10 min read -
Agent EngineeringHow cheap models changed multi-agent economics
Orchestrator-executor just became the smart default for production agents: an expensive model plans, cheap models execute, and cost per completed task decides the roster. Laurie Voss August 7, 2026 8 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.