The Evaluator newsletter

The agent feedback loop, in your inbox.

New playbooks, field notes, and frameworks for building reliable AI agents.

Guides

Go deep, chapter by chapter.

Long-form handbooks you can read end to end, or drop into at the chapter you need.

Browse all resources →
Handbook

What are AI agents? Architecture, tools & how they work

A production workflow for evaluating AI agents: define success, build datasets, trace completely, choose evaluators, run experiments, hold regression cases, and gate releases.

  1. 01 AI agent frameworks compared: LangGraph, CrewAI, AutoGen, and more
  2. 02 Agent observability: how to trace, debug, and improve AI agents
  3. 03 How to evaluate AI agents: a production workflow
  4. 04 Agent evaluation metrics: how to measure whether an agent works
Start the guide → 4 chapters
Guide

The definitive guide to LLM evaluations

LLM evaluation: Get from pre-production to deployment with our definitive guide to LLM evaluation. Includes LLM eval types, use cases, templates and tips for continuous improvement.

  1. 01 LLM evaluation metrics: correctness, groundedness, RAG & agent scores
  2. 02 Pre-production LLM evaluation: datasets, synthetic data & benchmarks
  3. 03 CI/CD for LLM apps: experiments, regression tests & release gates
  4. 04 Production LLM evaluation: guardrails, online evals & monitoring
Start the guide → 5 chapters
Videos & Talks

Demos, workshops & conference talks.

Watch on YouTube →
Rise of the AI Engineer

An agent got the right answer the wrong way | Michael Grinich, WorkOS

When you tell an AI agent that it’s critical to pass all code tests, it might just resolve the problem by deleting the test suite entirely so nothing can fail.

Don't ship vibes.

Trace, evaluate, and continuously improve your agents — built on open source & open standards.