Practical guides, field notes & frameworks for reliable AI agents.
Code-first for engineers. Quality frameworks for product managers. Operating models for leaders.
Agent harnesses: How to trace, evaluate, and improve AI agents
Learn how agent harnesses use tracing and evaluations to make AI agents observable, testable, safer, and easier to improve in production.
The definitive guide to LLM evaluations
LLM evaluation: Get from pre-production to deployment with our definitive guide to LLM evaluation. Includes LLM eval types, use cases, templates and…
AI agent evaluation: An agent-native framework
Learn how to evaluate AI agents across outcomes, trajectories, decisions, and repeated-run reliability using traces, checks, and LLM judges.
Why do you need evals for AI agents?
Agent failures hide inside correct-looking outputs. Learn why evals are the only mechanism that catches them, and how to wire them into…
Latest field notes & frameworks.
Cost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models
Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model…
How to measure human-LLM judge alignment
No single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human…
How OpenAI uses human feedback to evaluate and improve LLMs
At ChatGPT scale, user frustration arrives as support tickets, ratings, social posts, and corrections buried inside conversations. OpenAI built a feedback system…
Real teams, shipping AI.
How TheFork Leverages Online Evals To Boost Conversions with Arize AX on AWS
TheFork is one of Europe’s leading restaurant discovery and booking platforms, connecting millions of diners with tens of thousands of restaurants across…
Agents in the Wild: Priceline’s Journey into Evaluating Voice Applications
Join us for an inside look at how Priceline is evolving its AI-powered travel assistant, Penny, from text-based interactions to voice-enabled experiences. We’ll…
How Handshake Deployed and Scaled 15+ LLM Use Cases In Under Six Months — With Evals From Day One
Handshake is the largest early-career network, specializing in connecting students and new grads with employers and career centers. It’s also an engineering…
Demos, workshops & conference talks.
An agent got the right answer the wrong way | Michael Grinich, WorkOS
When you tell an AI agent that it’s critical to pass all code tests, it might just resolve the problem by deleting the test suite entirely so nothing can fail.
Don't ship vibes.
Trace, evaluate, and continuously improve your agents — built on open source & open standards.