-
Agent EngineeringAI agent regression testing with Agent Experiments in Arize AX
A cancellation-policy fix raised action safety and dropped average task completion from 0.89 to 0.72. This walkthrough shows how to regression-test agent changes with Agent Experiments in Arize… Nancy Chauhan Fuad Ali September 14, 2026 9 min read -
Agent EngineeringCode mode: Why your agent should code
Code mode gives an agent a sandbox instead of a longer tool list. Here is why it fixes the too-many-tools problem, what it costs you in sandboxing and… Mikyo King September 9, 2026 11 min read -
Agent EngineeringHow I cut coding agent costs with model and harness routing
By routing planning, exploration, implementation, and review to different models, I reduced one recurring coding-agent workflow from roughly $100 to $15-$20 per run. Arda Hoke September 8, 2026 9 min read -
Agent EngineeringHow Coinbase Wallet built an agent-first product development lifecycle
By redesigning planning, validation, and risk review around AI agents, Coinbase Wallet dramatically shortened the path from product idea to working software. Sara Verdi September 2, 2026 7 min read -
Agent EngineeringAgent cost management is about more than the model
Every LLM call your application makes costs money, and agentic applications make a lot of LLM calls. Arize AX now ships a Cost Agent that reads traces, ranks… Laurie Voss September 1, 2026 8 min read -
Agent EngineeringHow Signal found two hidden retry loops in our production agent Alyx
We ran Signal on Alyx, the AI engineering agent built into Arize AX. It surfaced a duplicate task-state loop and a 43-call dataset retry that appeared as valid… Nancy Chauhan August 27, 2026 8 min read -
Agent EngineeringWhy better models don’t fix every agent failure: Lessons from OpenAI
In this installment of Rise of the AI Engineer, Stuart Sy from OpenAI, explains why the bottleneck has moved off the model and onto context, evals, and observability. Sara Verdi August 25, 2026 9 min read -
Agent EngineeringHow Uber evaluates AI agents at production scale
A background comment about pizza exposed a failure that Uber’s offline evaluations had missed. The incident helped reveal what production AI agent evaluation actually requires: automatic tracing, living… Sara Verdi August 14, 2026 14 min read -
Agent EngineeringEvaluation-driven development: How to move AI agents from pilot to production
Learn how evaluation-driven development, agent harnesses, AI observability, guardrails, and cost-per-outcome metrics move AI agents from pilot to production. Sara Verdi August 13, 2026 19 min read -
Agent EngineeringHow cheap models changed multi-agent economics
Orchestrator-executor just became the smart default for production agents: an expensive model plans, cheap models execute, and cost per completed task decides the roster. Laurie Voss August 7, 2026 8 min read -
Agent EngineeringHamel Husain explains why AI evals fail before the evaluation begins
Hamel Husain explains why ambiguous inputs, generic metrics, and disconnected review workflows can make AI evaluations misleading, and how developers can build a better process around real production… Sara Verdi July 30, 2026 7 min read -
Agent EngineeringHow to improve agent skills with tracing and evals
A skill cut agent costs by 44% and latency by 56%, but also reduced answer completeness. Here’s how tracing, evals, and a long-running agent exposed and corrected the… Yusuf Cattaneo July 28, 2026 8 min read -
Agent EngineeringTips from Anthropic on building agent evals you can trust
Learn how to build trustworthy AI agent evals using regression tests, capability evals, production traces, LLM judges, and reproducible environments. Sara Verdi July 28, 2026 17 min read -
Agent EngineeringHow to write effective AI agent skills: 6 data-backed practices
Three recent studies show what actually makes an AI agent skill effective: human expertise, compact procedures, tight routing, harness-specific testing, and eval-gated changes—not longer Markdown. Laurie Voss July 24, 2026 11 min read -
Agent EngineeringInside Cursor’s agent factory: how it verifies AI-written code
As background agents take on more implementation work, Cursor is rebuilding the software development lifecycle around risk scores, developer-like environments, video evidence, and review systems that learn from… Sara Verdi July 20, 2026 10 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.