Guides — page 2.
What “self-hosted” actually means in AI observability
Learn what self-hosted AI observability means. Compare SaaS, hybrid, open-source, private, and air-gapped deployments with a vendor checklist.
Read the guide
AI observability pricing: how traces, spans, scores, and seats change your bill
Compare AI observability pricing across LangSmith, Langfuse, Braintrust, Datadog, and Arize. See what each vendor meters, how evals…
Read the guide
7 best AI agent evaluation platforms compared for 2026
The best LLM evaluation platform is the one that can evaluate the units your application actually produces, run…
Read the guide
What is an agent observability platform?
Learn why APM fails for autonomous agents and how an agent observability platform uses trace-centric analysis and evaluation…
Read the guide
What are AI agents? Architecture, tools & how they work
Learn what AI agents are, how they work, and how to build them. Explore agent architecture, tools, memory,…
Read more
LLM-as-a-Judge: When should you use it?
LLM judges fit a narrow window. Learn the specific conditions that qualify a task, three prerequisites, and when…
Read the guide
AI agent tracing and evaluation: The complete developer guide
Learn how to trace and evaluate AI agents across spans, trajectories, and sessions. Build reliable evals with OpenTelemetry,…
Read the guide
What is an AI product manager? A guide to the role and skills (2026)
What separates an AI product manager from a traditional PM and how to become an one, with tips…
Read the guide
What is an agent harness? Architecture, controls, and evaluation
Learn how agent harnesses use tracing and evaluations to make AI agents observable, testable, safer, and easier to…
Read the guide
The definitive guide to LLM evaluations
LLM evaluation: Get from pre-production to deployment with our definitive guide to LLM evaluation. Includes LLM eval types,…
Read the guide
What is AI engineering?
AI engineering builds applications on top of models; it treats the model as a component rather than the…
Read the guide
LLM evaluation costs: Understanding hidden costs & budget models
Learn what drives LLM evaluation costs across offline tests and production, from judge tokens and trace volume to…
Read the guideDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.