-
LLM EvaluationHamel Husain explains why AI evals fail before the evaluation begins
Hamel Husain explains why ambiguous inputs, generic metrics, and disconnected review workflows can make AI evaluations misleading, and how developers can build a better process around real production… Sara Verdi July 30, 2026 7 min read -
LLM EvaluationHow to measure human-LLM judge alignment
No single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human agreement, and diagnose errors with precision,… Elizabeth Hutton July 22, 2026 15 min read -
LLM EvaluationHow do you make an LLM, anyway? Microsoft just published a textbook.
Microsoft published a 109-page technical report on MAI-Thinking-1. Here’s the abbreviated version of how a modern lab actually trains a frontier reasoning model — from scraping the web… Laurie Voss July 13, 2026 11 min read -
LLM EvaluationEvals in CI: How to write your LLM evals as tests with Arize Phoenix
If you're struggling to get started with evals, you're not alone. This post explains how to write LLM evals as ordinary tests in CI with Phoenix, pytest, and… Mikyo King July 7, 2026 17 min read -
LLM EvaluationTrace and evaluate TrueFoundry AI Gateway traffic in Arize AX
Learn how TrueFoundry AI Gateway exports OpenTelemetry traces to Arize AX so teams can trace, evaluate, and monitor production LLM and agent traffic without embedding a vendor SDK… Aaron Winston June 29, 2026 7 min read -
LLM EvaluationBuilding the AI factory for self-improving agents: What’s new in Arize AX
Arize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. Jason Lopatecki Aparna Dhinakaran June 4, 2026 8 min read -
LLM EvaluationHow to ship a local LLM that matches frontier LLMs with evals and prompt engineering
Most production AI features don't need a frontier model. Here's how capability evals and prompt engineering can help ship a local SLM that matches frontier-model quality with lower… RL Nabors May 26, 2026 15 min read -
LLM EvaluationHow to build LLM-as-a-Judge evaluators that hold up in production
Learn how to design, calibrate, and run LLM-as-a-judge evaluators with fixed labels, human agreement checks, trace context, and Phoenix Evals. Aaron Winston May 21, 2026 22 min read -
LLM EvaluationHow We Used Evals (and an AI Agent) to Iteratively Improve an AI Newsletter Generator
We love building little AI-powered tools that accelerate our workflows. One we built recently is a tool that takes our recent tweets and uses Claude to create a… Laurie Voss March 10, 2026 10 min read -
LLM Evaluation14 best AI agent observability tools in 2026: A practical comparison
Compare 14 AI agent observability tools for tracing, evaluations, OpenTelemetry, self-hosting, pricing, and production monitoring. Updated July 2026. Aryan Kargwal February 27, 2026 29 min read -
LLM EvaluationMastering Production RAG with Google ADK and Arize AX for Enterprise Knowledge Systems
Introduction Retrieval Augmented Generation (RAG) has become the cornerstone of enterprise AI, yet most organizations struggle with a critical challenge: building RAG systems that work reliably in production.… Richard Young February 23, 2026 11 min read -
LLM EvaluationNew In Arize AX: January 2026 Updates
Arize AX pushed out a lot of new updates in January 2026. From improved evaluator hub to custom prompt release labels, here are some highlights. Evaluator Hub: Reusable… Sanjana Yeddula February 2, 2026 8 min read -
LLM EvaluationHow TheFork Leverages Online Evals To Boost Conversions with Arize AX on AWS
TheFork is one of Europe’s leading restaurant discovery and booking platforms, connecting millions of diners with tens of thousands of restaurants across major cities. The company’s marketplace spans… Yesmine Rouis Natalia Skaczkowska-Drabczyk December 9, 2025 4 min read -
LLM EvaluationGEPA vs Prompt Learning: Benchmarking Different Prompt Optimization Approaches
In June 2025, Andrej Karpathy introduced Software 3.0: the notion that software development is shifting from programming through code to prompting through natural language. When building programs, the… Priyan Jindal November 17, 2025 11 min read -
LLM EvaluationServiceNow’s Tara Bogavelli on AgentArch: Benchmarking AI Agents for Enterprise Workflows
In our latest AI research paper reading, we hosted Tara Bogavelli, Machine Learning Engineer at ServiceNow, to discuss her team’s recent work on AgentArch, a new benchmark designed… Julian Reeves October 24, 2025 4 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.