Latest Resources

Practical guides, field notes & frameworks for reliable AI agents.

See all →

Code-first for engineers. Quality frameworks for product managers. Operating models for leaders.

Guide

Best LLM and agent evaluation platforms: A comparison

The best LLM evaluation platform is the one that can evaluate the units your application actually produces, run consistent quality criteria before and after deployment, align automated scores with human judgment, and turn failures into reproducible…

Chris Cooning 34 min read Jul 2026
The Evaluator newsletter

The agent feedback loop, in your inbox.

New playbooks, field notes, and frameworks for building reliable AI agents.

Videos & Talks

Demos, workshops & conference talks.

Watch on YouTube →
Rise of the AI Engineer

An agent got the right answer the wrong way | Michael Grinich, WorkOS

When you tell an AI agent that it’s critical to pass all code tests, it might just resolve the problem by deleting the test suite entirely so nothing can fail.

Don't ship vibes.

Trace, evaluate, and continuously improve your agents — built on open source & open standards.