Resource Hub

Showing 1–12 of 830 resources (page 1 of 70)

How Uber evaluates AI agents at production scale
Blog

How Uber evaluates AI agents at production scale

A background comment about pizza exposed a failure that Uber’s offline evaluations had missed. The incident helped reveal what production AI agent evaluation actually requires: automatic tracing, living datasets, shared ownership, and a direct connection to product decisions.

AI model lifecycle management: 7 stages, controls, and tools

AI model lifecycle management: 7 stages, controls, and tools

The seven stages of AI model lifecycle management, what to version at each gate, which tools own which record, and what changes for LLM apps and AI agents.

Arize and Dynatrace: Making the World’s AI Work
Blog

Arize and Dynatrace: Making the World’s AI Work

Today we are announcing the signing of a definitive agreement for the acquisition of Arize by Dynatrace to accelerate our mission to make the world's AI work.

AI agent guardrails vs. evals: How to build more reliable agent systems

AI agent guardrails vs. evals: How to build more reliable agent systems

Guardrails constrain what an agent can do in code; evals judge whether it performed well. Learn how both layers—and the harness around them—make long-running AI agents reliable.

Evaluation-driven development: How to move AI agents from pilot to production
Blog

Evaluation-driven development: How to move AI agents from pilot to production

Learn how evaluation-driven development, agent harnesses, AI observability, guardrails, and cost-per-outcome metrics move AI agents from pilot to production.

Crew Studio launches with native Arize AX tracing and evaluation

Crew Studio launches with native Arize AX tracing and evaluation

Through a native Arize AX integration, teams can send traces from Crew Studio to Arize from the first run without custom instrumentation—then inspect behavior, evaluate quality, and test fixes before redeploying.

AI agent debugging tools: 9 platforms compared for production in 2026

AI agent debugging tools: 9 platforms compared for production in 2026

What separates a trace viewer from a production debugging system? Compare 9 tools across failure discovery, diagnosis, regression testing, and remediation.

You chose the best model. Why is your agent still failing?
Blog

You chose the best model. Why is your agent still failing?

Public benchmarks can show how a model performs in general. Production reliability depends on the context and harness around it, which only your team can evaluate against its own data, workflows, and users.

AI agent testing: 7 failures traditional software tests miss

AI agent testing: 7 failures traditional software tests miss

Traditional tests can confirm that an AI agent request completed. Learn the seven task, tool, trajectory, and session failures that agent evals expose.

Harness engineering: how to build reliable AI agents

Harness engineering: how to build reliable AI agents

Harness engineering is how you govern a production agent runtime: task contracts, completion gates, deterministic authority, durable checkpoints, and side-effect-aware recovery.

Arize AX adds native support for OpenTelemetry GenAI semantic conventions

Arize AX adds native support for OpenTelemetry GenAI semantic conventions

Arize AX now normalizes OpenTelemetry GenAI semantic conventions into first-class AI traces, unlocking evaluations, token and cost visibility, and easier debugging.

What “self-hosted” actually means in AI observability

What “self-hosted” actually means in AI observability

Learn what self-hosted AI observability means. Compare SaaS, hybrid, open-source, private, and air-gapped deployments with a vendor checklist.

No results found. Try a different filter or search term.