-
LLM EvalsHamel Husain explains why AI evals fail before the evaluation begins
Hamel Husain explains why ambiguous inputs, generic metrics, and disconnected review workflows can make AI evaluations misleading, and how developers can build a better process around real production… Sara Verdi July 30, 2026 7 min read -
LLM EvalsHow to write effective AI agent skills: 6 data-backed practices
Three recent studies show what actually makes an AI agent skill effective: human expertise, compact procedures, tight routing, harness-specific testing, and eval-gated changes—not longer Markdown. Laurie Voss July 24, 2026 11 min read -
LLM EvalsHow to measure human-LLM judge alignment
No single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human agreement, and diagnose errors with precision,… Elizabeth Hutton July 22, 2026 16 min read -
LLM EvalsHow do you make an LLM, anyway? Microsoft just published a textbook.
Microsoft published a 109-page technical report on MAI-Thinking-1. Here’s the abbreviated version of how a modern lab actually trains a frontier reasoning model — from scraping the web… Laurie Voss July 13, 2026 11 min read -
LLM EvalsAI evals are a data science problem: What most teams get wrong
Hamel Husain explains why the best AI teams treat LLM judges like classifiers, not dashboards. Sara Verdi June 30, 2026 10 min read -
LLM EvalsMeet PXI: the AI engineering agent inside Phoenix
An AI engineering agent built into Phoenix. It works like a coding agent, just point it at your telemetry instead of a source tree. Mikyo King Roger Yang Nancy Chauhan Anthony Powell June 18, 2026 17 min read -
LLM EvalsHow to ship a local LLM that matches frontier LLMs with evals and prompt engineering
Most production AI features don't need a frontier model. Here's how capability evals and prompt engineering can help ship a local SLM that matches frontier-model quality with lower… RL Nabors May 26, 2026 15 min read -
LLM EvalsFrom First Eval to Autonomous AI Ops: A Maturity Model for AI Evaluation
Every team runs evals. Almost none have an evaluation practice. The difference is the gap between a one-off notebook and a system that continuously assesses, alerts, and acts… Cam Young April 3, 2026 6 min read -
LLM Evals14 best AI agent observability tools in 2026: A practical comparison
Compare 14 AI agent observability tools for tracing, evaluations, OpenTelemetry, self-hosting, pricing, and production monitoring. Updated July 2026. Aryan Kargwal February 27, 2026 29 min read -
LLM EvalsAI Agent Debugging: Four Lessons from Shipping Alyx to Production
Building AI systems that actually work in production is harder than it sounds. Not demo-ware, not “it worked once in a notebook.” Real systems that keep working after… Laurie Voss February 25, 2026 22 min read -
LLM EvalsServiceNow’s Tara Bogavelli on AgentArch: Benchmarking AI Agents for Enterprise Workflows
In our latest AI research paper reading, we hosted Tara Bogavelli, Machine Learning Engineer at ServiceNow, to discuss her team’s recent work on AgentArch, a new benchmark designed… Julian Reeves October 24, 2025 4 min read -
LLM EvalsWhat Are the Top LLM Evaluation Tools?
AI agents and real-world applications of generative AI are debuting at an incredible clip this year, narrowing the time from AI research paper to industry application and propelling… David Burch October 23, 2025 2 min read -
LLM EvalsShould I Use the Same LLM for My Eval as My Agent? Testing Self-Evaluation Bias
Thanks to Aparna Dhinakaran and Elizabeth Hutton for their contributions to this piece. When building and testing AI agents, one practical question that arises is whether to use… Sanjana Yeddula October 8, 2025 10 min read -
LLM EvalsTesting Binary vs Score Evals on the Latest Models
Thanks to Hamel Husain and Eugene Yan for reviewing this piece Evals are becoming the predominant approach for how AI engineers systematically evaluate the quality of the LLM… Aparna Dhinakaran Sri Chavali September 24, 2025 10 min read -
LLM EvalsAI Evals Maven Course Homework: the Recipe Bot Workflow
AI Evals for Engineers & PMs is a popular, hands‑on Maven course led by Hamel Husain and Shreya Shankar. The course’s goal is simple: “teach a systematic workflow… Sri Chavali September 3, 2025 9 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.