Sara Verdi
-
Agent Engineering3 production patterns for AI agents and how to evaluate each one
A local coding agent, an in-app customer assistant, and an AI SRE triaging production logs may all use the same model class—but not the same harness, eval plan,… Sara Verdi July 10, 2026 9 min read -
Agent EvaluationTrace before you migrate: Measuring Kubernetes bottlenecks in AI agent sandboxes
Kubernetes is strong for long-lived services, but it is often a poor default for short-lived agent sandboxes. Trace sandbox creation, tool execution, eval latency, and full trajectory time… Sara Verdi July 9, 2026 7 min read -
Agent EvaluationHow to evaluate AI agents, avoid reward hacking, and build better specs
Agent evals are repeatable tests that score whether AI agents completed a task correctly. Learn how to design rubrics, test suites, and trace-based evals that catch failures and… Sara Verdi July 2, 2026 9 min read -
AI EvaluationAI evals are a data science problem: What most teams get wrong
Hamel Husain explains why the best AI teams treat LLM judges like classifiers, not dashboards. Sara Verdi June 30, 2026 10 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.