-
Agent EvaluationAI agent regression testing with Agent Experiments in Arize AX
A cancellation-policy fix raised action safety and dropped average task completion from 0.89 to 0.72. This walkthrough shows how to regression-test agent changes with Agent Experiments in Arize… Nancy Chauhan Fuad Ali September 14, 2026 9 min read -
Agent EvaluationHow I cut coding agent costs with model and harness routing
By routing planning, exploration, implementation, and review to different models, I reduced one recurring coding-agent workflow from roughly $100 to $15-$20 per run. Arda Hoke September 8, 2026 9 min read -
Agent EvaluationArize Phoenix has a built-in MCP server that lets your agents query traces with SQL
Read-only SQL and code mode let coding agents answer questions across your traces without paging thousands of spans through model context. Nancy Chauhan Roger Yang Mora Vigo Malusardi August 27, 2026 10 min read -
Agent EvaluationWhy better models don’t fix every agent failure: Lessons from OpenAI
In this installment of Rise of the AI Engineer, Stuart Sy from OpenAI, explains why the bottleneck has moved off the model and onto context, evals, and observability. Sara Verdi August 25, 2026 9 min read -
Agent EvaluationA skill is just an agent. So measure your changes.
A skill is just another AI agent: a prompt plus a harness that runs it. That means you can trace it, eval it, and prove a change made… Jim Bennett August 24, 2026 9 min read -
Agent EvaluationIs your coding agent uploading all your code?
After Grok Build was caught uploading entire Git repos, we read the privacy docs for Claude Code, Codex, Cursor, GitHub Copilot, and Grok Build to compare what code… Laurie Voss August 20, 2026 7 min read -
Agent EvaluationWhere agent evals are going: Agent-as-a-Judge
Agents changed what failure looks like, and the evaluation layer has to change with them. Why agent-as-a-judge is moving from research paper to production eval stack. Laurie Voss August 19, 2026 8 min read -
Agent EvaluationHow Uber evaluates AI agents at production scale
A background comment about pizza exposed a failure that Uber’s offline evaluations had missed. The incident helped reveal what production AI agent evaluation actually requires: automatic tracing, living… Sara Verdi August 14, 2026 14 min read -
Agent EvaluationHow cheap models changed multi-agent economics
Orchestrator-executor just became the smart default for production agents: an expensive model plans, cheap models execute, and cost per completed task decides the roster. Laurie Voss August 7, 2026 8 min read -
Agent EvaluationHamel Husain explains why AI evals fail before the evaluation begins
Hamel Husain explains why ambiguous inputs, generic metrics, and disconnected review workflows can make AI evaluations misleading, and how developers can build a better process around real production… Sara Verdi July 30, 2026 7 min read -
Agent EvaluationHow to improve agent skills with tracing and evals
A skill cut agent costs by 44% and latency by 56%, but also reduced answer completeness. Here’s how tracing, evals, and a long-running agent exposed and corrected the… Yusuf Cattaneo July 28, 2026 8 min read -
Agent EvaluationTips from Anthropic on building agent evals you can trust
Learn how to build trustworthy AI agent evals using regression tests, capability evals, production traces, LLM judges, and reproducible environments. Sara Verdi July 28, 2026 17 min read -
Agent EvaluationHow to write effective AI agent skills: 6 data-backed practices
Three recent studies show what actually makes an AI agent skill effective: human expertise, compact procedures, tight routing, harness-specific testing, and eval-gated changes—not longer Markdown. Laurie Voss July 24, 2026 11 min read -
Agent EvaluationCost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models
Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model evaluation and routing. Laurie Voss July 23, 2026 16 min read -
Agent EvaluationKiro CLI observability: trace and evaluate agent changes with Arize Skills
Use Arize Skills with Kiro CLI to trace coding-agent changes, build datasets from failures, run experiments, and validate prompts before shipping. Richard Young July 15, 2026 11 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.