Practical guides, field notes & frameworks for reliable AI agents.
Code-first for engineers. Quality frameworks for product managers. Operating models for leaders.
AI observability pricing: how traces, spans, scores, and seats change your bill
Compare AI observability pricing across LangSmith, Langfuse, Braintrust, Datadog, and Arize. See what each vendor meters, how evals affect cost, and what to model before you buy.
LLM & agent evaluation platforms: 2026 comparison
The best LLM evaluation platform is the one that can evaluate the units your application actually produces, run consistent quality criteria before…
What is an agent observability platform?
Learn why APM fails for autonomous agents and how an agent observability platform uses trace-centric analysis and evaluation workflows to prevent failures.
AI agent tracing and evaluation: The complete developer guide
Learn how to trace and evaluate AI agents across spans, trajectories, and sessions. Build reliable evals with OpenTelemetry, OpenInference, and Arize AX.
Latest field notes & frameworks.
Demystifying the EU AI Act for AI product and engineering teams
An engineering guide to turning EU AI Act principles into traces, evaluations, annotations, and release evidence product and engineering teams can actually…
How cheap models changed multi-agent economics
Orchestrator-executor just became the smart default for production agents: an expensive model plans, cheap models execute, and cost per completed task decides…
AI agent observability: Why production systems need a reasoning layer
Traditional APM can collect every span and still leave developers guessing about intent, causality, and drift. As agents multiply, the observability stack…
Real teams, shipping AI.
How Booking.com scales AI observability with Arize
How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and…
How TheFork Leverages Online Evals To Boost Conversions with Arize AX on AWS
TheFork is one of Europe’s leading restaurant discovery and booking platforms, connecting millions of diners with tens of thousands of restaurants across…
How Handshake Deployed and Scaled 15+ LLM Use Cases In Under Six Months — With Evals From Day One
Handshake is the largest early-career network, specializing in connecting students and new grads with employers and career centers. It’s also an engineering…
Demos, workshops & conference talks.
An agent got the right answer the wrong way | Michael Grinich, WorkOS
When you tell an AI agent that it’s critical to pass all code tests, it might just resolve the problem by deleting the test suite entirely so nothing can fail.
Don't ship vibes.
Trace, evaluate, and continuously improve your agents — built on open source & open standards.