Evaluations Quickstart
Get started running evaluations to measure how your model performs.
Align LLM Evals with Human Judgment
Iteratively refine a custom LLM-as-a-Judge evaluator against human-annotated ground truth.
Why Public Benchmarks Lie: Building Your Own Eval Harness
Build your own eval harness instead of trusting public benchmarks, via an email-extraction service.
Trace-Level Evaluations for a Recommendation Agent
Run trace-level evaluations on individual requests to a recommendation agent.
Session-Level Evaluations for an AI Tutor
Run multi-dimensional session-level evaluations on multi-turn AI tutor conversations.
Evaluating RAG Retrieval Quality and Correctness
Trace, evaluate, diagnose, and improve retrieval quality and correctness in a RAG application.
Evaluating Agentic RAG Using Arize AX and Couchbase
Build and evaluate an agentic RAG application on a Couchbase vector store.
Evaluate a Math Problem-Solving Agent Using Ragas
Create and evaluate a math problem-solving agent using Ragas and Arize AX.
Pydantic Evals
Evaluate a question-answering task with Pydantic Evals and log results to Arize AX.
Tracing and Evaluating Voice Applications
Trace OpenAI Realtime voice agents and run tone evaluation on captured audio.
Audio Transcription and Evaluation with Gemini Flash
Transcribe and evaluate audio with Gemini Flash, traced in Arize AX.
More Guides
Span-level evaluator examples for hallucination, relevance, toxicity, SQL, tool calling, and more.