Everything we’ve published.
-
Agent EngineeringAI agent regression testing with Agent Experiments in Arize AX
A cancellation-policy fix raised action safety and dropped average task completion from 0.89 to 0.72. This walkthrough shows how to regression-test agent changes with Agent Experiments in Arize… Nancy Chauhan Fuad Ali September 14, 2026 9 min read -
Agent EngineeringCode mode: Why your agent should code
Code mode gives an agent a sandbox instead of a longer tool list. Here is why it fixes the too-many-tools problem, what it costs you in sandboxing and… Mikyo King September 9, 2026 11 min read -
Agent EngineeringHow I cut coding agent costs with model and harness routing
By routing planning, exploration, implementation, and review to different models, I reduced one recurring coding-agent workflow from roughly $100 to $15-$20 per run. Arda Hoke September 8, 2026 9 min read -
Agent EngineeringHow Coinbase Wallet built an agent-first product development lifecycle
By redesigning planning, validation, and risk review around AI agents, Coinbase Wallet dramatically shortened the path from product idea to working software. Sara Verdi September 2, 2026 7 min read -
Agent EngineeringAgent cost management is about more than the model
Every LLM call your application makes costs money, and agentic applications make a lot of LLM calls. Arize AX now ships a Cost Agent that reads traces, ranks… Laurie Voss September 1, 2026 8 min read -
AI EngineeringHow to reduce LLM costs without sacrificing quality
Trace where your AI budget goes, identify the requests that do not justify their spend, and validate lower-cost alternatives with Arize AX. Nancy Chauhan August 31, 2026 11 min read -
Agent EngineeringHow Signal found two hidden retry loops in our production agent Alyx
We ran Signal on Alyx, the AI engineering agent built into Arize AX. It surfaced a duplicate task-state loop and a 43-call dataset retry that appeared as valid… Nancy Chauhan August 27, 2026 8 min read -
Agent EvaluationArize Phoenix has a built-in MCP server that lets your agents query traces with SQL
Read-only SQL and code mode let coding agents answer questions across your traces without paging thousands of spans through model context. Nancy Chauhan Roger Yang Mora Vigo Malusardi August 27, 2026 10 min read -
Agent EngineeringWhy better models don’t fix every agent failure: Lessons from OpenAI
In this installment of Rise of the AI Engineer, Stuart Sy from OpenAI, explains why the bottleneck has moved off the model and onto context, evals, and observability. Sara Verdi August 25, 2026 9 min read -
Agent EvaluationA skill is just an agent. So measure your changes.
A skill is just another AI agent: a prompt plus a harness that runs it. That means you can trace it, eval it, and prove a change made… Jim Bennett August 24, 2026 9 min read -
Agent EvaluationIs your coding agent uploading all your code?
After Grok Build was caught uploading entire Git repos, we read the privacy docs for Claude Code, Codex, Cursor, GitHub Copilot, and Grok Build to compare what code… Laurie Voss August 20, 2026 7 min read -
Agent EvaluationWhere agent evals are going: Agent-as-a-Judge
Agents changed what failure looks like, and the evaluation layer has to change with them. Why agent-as-a-judge is moving from research paper to production eval stack. Laurie Voss August 19, 2026 8 min read -
Agent EngineeringHow Uber evaluates AI agents at production scale
A background comment about pizza exposed a failure that Uber’s offline evaluations had missed. The incident helped reveal what production AI agent evaluation actually requires: automatic tracing, living… Sara Verdi August 14, 2026 14 min read -
CompanyArize and Dynatrace: Making the World’s AI Work
Today we are announcing the signing of a definitive agreement for the acquisition of Arize by Dynatrace to accelerate our mission to make the world's AI work. Jason Lopatecki August 13, 2026 5 min read -
AgentsAI agent guardrails vs. evals: How to build more reliable agent systems
Guardrails constrain what an agent can do in code; evals judge whether it performed well. Learn how both layers—and the harness around them—make long-running AI agents reliable. Aaron Winston August 13, 2026 9 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.