Arize Phoenix has a built-in MCP server that lets your agents query traces with SQL
Read-only SQL and code mode let coding agents answer questions across your traces without paging thousands of spans through model context.
10 min read
Why better models don’t fix every agent failure: Lessons from OpenAI
In this installment of Rise of the AI Engineer, Stuart Sy from OpenAI, explains why the bottleneck has moved off the model and onto context, evals, and observability.
9 min read
A skill is just an agent. So measure your changes.
A skill is just another AI agent: a prompt plus a harness that runs it. That means you can trace it, eval it, and prove a change made…
9 min readDemos, workshops, and conference talks.
See all on YouTube
Why AI Agents Fail at Tasks They Already Completed | Ivan Burazin, Daytona
Ivan Burazin (CEO & Co-founder of Daytona) breaks down why the model is rarely the bottleneck; tool access, broken credential boundaries, and flawed harness design are.
InsightsGet the latest on AI & Observability
Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:
Real teams, shipping AI.
See allHow Booking.com scales AI observability with Arize
How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and…
Read moreHow LG Uplus is building better AI customer service agents with evaluation-driven development
How LG Uplus uses Arize AX to build evaluation-driven AI contact center agents — combining production traces, user feedback, and domain expertise to continuously improve customer service for…
Read moreHow Tripadvisor is building the AI product development lifecycle for agentic travel
Tripadvisor VP of Data and AI Rahul Todkar on building a production AI lifecycle for traditional ML and agentic travel products with unified observability, traces, evals, and governance…
Read moreCatch up on everything you missed.
See all-
Agent EvaluationIs your coding agent uploading all your code?
After Grok Build was caught uploading entire Git repos, we read the privacy docs for Claude Code, Codex, Cursor, GitHub Copilot, and Grok Build to compare what code… Laurie Voss August 20, 2026 7 min read -
Agent EvaluationWhere agent evals are going: Agent-as-a-Judge
Agents changed what failure looks like, and the evaluation layer has to change with them. Why agent-as-a-judge is moving from research paper to production eval stack. Laurie Voss August 19, 2026 8 min read -
Agent EngineeringHow Uber evaluates AI agents at production scale
A background comment about pizza exposed a failure that Uber’s offline evaluations had missed. The incident helped reveal what production AI agent evaluation actually requires: automatic tracing, living… Sara Verdi August 14, 2026 14 min read -
CompanyArize and Dynatrace: Making the World’s AI Work
Today we are announcing the signing of a definitive agreement for the acquisition of Arize by Dynatrace to accelerate our mission to make the world's AI work. Jason Lopatecki August 13, 2026 5 min read -
AgentsAI agent guardrails vs. evals: How to build more reliable agent systems
Guardrails constrain what an agent can do in code; evals judge whether it performed well. Learn how both layers—and the harness around them—make long-running AI agents reliable. Aaron Winston August 13, 2026 9 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.