Nancy Chauhan
-
Agent EngineeringPrompt caching benchmark: high cache reuse doesn’t always mean lower cost
We benchmarked prompt caching across DeepSeek, GLM, GPT, and Claude using Harbor evals and Phoenix traces, so we could compare cache reuse, estimated cost, and latency on the… Nancy Chauhan October 2, 2026 7 min read -
Agent EngineeringAI agent regression testing with Agent Experiments in Arize AX
A cancellation-policy fix raised action safety and dropped average task completion from 0.89 to 0.72. This walkthrough shows how to regression-test agent changes with Agent Experiments in Arize… Nancy Chauhan Fuad Ali September 14, 2026 9 min read -
AI EngineeringHow to reduce LLM costs without sacrificing quality
Trace where your AI budget goes, identify the requests that do not justify their spend, and validate lower-cost alternatives with Arize AX. Nancy Chauhan August 31, 2026 11 min read -
Agent EngineeringHow Signal found two hidden retry loops in our production agent Alyx
We ran Signal on Alyx, the AI engineering agent built into Arize AX. It surfaced a duplicate task-state loop and a 43-call dataset retry that appeared as valid… Nancy Chauhan August 27, 2026 8 min read -
Agent EvaluationArize Phoenix has a built-in MCP server that lets your agents query traces with SQL
Read-only SQL and code mode let coding agents answer questions across your traces without paging thousands of spans through model context. Nancy Chauhan Roger Yang Mora Vigo Malusardi August 27, 2026 10 min read -
AI EngineeringHow to debug production AI agents with Signal in Arize AX
Learn how Arize Signal turns production traces into ranked issues, proposed fixes, regression datasets, and reviewable pull requests for AI agents. Nancy Chauhan August 4, 2026 16 min read -
Agent EngineeringMeet PXI: the AI engineering agent inside Phoenix
An AI engineering agent built into Phoenix. It works like a coding agent, just point it at your telemetry instead of a source tree. Mikyo King Roger Yang Nancy Chauhan Anthony Powell June 18, 2026 17 min read -
Agent EngineeringHow to detect credential theft in AI agent harness traces
In May 2026, a malicious version of a popular VS Code extension spent 18 minutes in the marketplace before anyone caught it. In that time it ran on… Nancy Chauhan June 9, 2026 14 min read -
Agent EvaluationPhoenix at 10,000 stars on GitHub: How an open source AI observability project grew by following its community
Phoenix crossed 10,000 GitHub stars. Here is how the open-source AI observability project grew from a Jupyter notebook extension into a community-shaped platform for traces, evals, OpenInference, and… RL Nabors Nancy Chauhan June 7, 2026 10 min read -
Agent EngineeringWhat we learned testing 7 models under the same agent harness
Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on… Nancy Chauhan May 20, 2026 10 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.