-
LLM As A JudgeHow to measure human-LLM judge alignment
No single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human agreement, and diagnose errors with precision,… Elizabeth Hutton July 22, 2026 16 min read -
LLM As A JudgeHow do you make an LLM, anyway? Microsoft just published a textbook.
Microsoft published a 109-page technical report on MAI-Thinking-1. Here’s the abbreviated version of how a modern lab actually trains a frontier reasoning model — from scraping the web… Laurie Voss July 13, 2026 11 min read -
LLM As A JudgeAI evals are a data science problem: What most teams get wrong
Hamel Husain explains why the best AI teams treat LLM judges like classifiers, not dashboards. Sara Verdi June 30, 2026 10 min read -
LLM As A JudgeBuilding the AI factory for self-improving agents: What’s new in Arize AX
Arize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. Jason Lopatecki Aparna Dhinakaran June 4, 2026 8 min read -
LLM As A JudgeHow to build LLM-as-a-Judge evaluators that hold up in production
Learn how to design, calibrate, and run LLM-as-a-judge evaluators with fixed labels, human agreement checks, trace context, and Phoenix Evals. Aaron Winston May 21, 2026 22 min read -
LLM As A JudgeShould I Use the Same LLM for My Eval as My Agent? Testing Self-Evaluation Bias
Thanks to Aparna Dhinakaran and Elizabeth Hutton for their contributions to this piece. When building and testing AI agents, one practical question that arises is whether to use… Sanjana Yeddula October 8, 2025 10 min read -
LLM As A JudgeEvidence-Based Prompting Strategies for LLM-as-a-Judge: Explanations and Chain-of-Thought
When LLMs are used as evaluators, two design choices often determine the quality and usefulness of their judgments: whether to require explanations for decisions, and whether to use… Sri Chavali Elizabeth Hutton Aparna Dhinakaran August 20, 2025 8 min read -
LLM As A JudgeLLM-as-a-Judge: Example of How To Build a Custom Evaluator Using a Benchmark Dataset
When To Build Custom Evaluators Arize-Phoenix ships with pre-built evaluators that are tested against benchmark datasets and tuned for repeatability. They’re a fast way to stand up rigorous… Sanjana Yeddula August 12, 2025 2 min read -
LLM As A JudgeLLMs as Judges: A Comprehensive Survey on LLM-Based Evaluation Methods
We discuss a major survey of the LLMs-as-Judges paradigm: “LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.” This paper systematically examines the LLMs-as-Judge framework across five dimensions: functionality,… Sarah Welsh December 23, 2024 3 min read -
LLM As A JudgeAgent-as-a-Judge: Evaluate Agents with Agents
This week we dive into a paper that presents the “Agent-as-a-Judge” framework, a new paradigm for evaluating agent systems. Where typical evaluation methods focus solely on outcomes or… Sarah Welsh November 22, 2024 3 min read -
LLM As A JudgeTracing and Evaluating LangGraph Agents
LangGraph is a powerful library designed for building stateful, multi-actor applications within large language models (LLMs). In this post, we’ll discuss how LangGraph’s traces can be ingested into… Greg Chase October 16, 2024 6 min read -
LLM As A JudgeBest Practices for Selecting the Right Model for LLM-as-a-Judge Evaluations
When building and scaling LLM-based applications, ensuring model performance is critical. One powerful method for evaluating that performance is using an LLM as a judge. This allows you… Samantha White September 30, 2024 5 min read -
LLM As A JudgeJudging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Introduction This week’s paper presents a comprehensive study of the performance of various LLMs acting as judges. The researchers leverage TriviaQA as a benchmark for assessing objective knowledge… Sarah Welsh August 16, 2024 40 min read -
LLM As A JudgeText To SQL: Evaluating SQL Generation with LLM as a Judge
Special shoutout to Manas Singh for collaborating with us on this research! One application of LLMs that has garnered headlines and significant investment surrounds their ability to generate… Aparna Dhinakaran Evan Jolley August 1, 2024 4 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.