What agent traces can tell you without an LLM judge
Some agent failures are already in the trace. tracelint proves them on Phoenix OpenInference spans and fails CI, without an LLM judge.
8 min read
How we benchmark AI agents and tools with Harbor and Arize Phoenix
See how Arize uses Harbor and Phoenix to benchmark AI agents, MCP servers, skills, and CLIs in reproducible sandboxes, then inspect traces and annotations.
5 min read
Prompt caching benchmark: high cache reuse doesn’t always mean lower cost
We benchmarked prompt caching across DeepSeek, GLM, GPT, and Claude using Harbor evals and Phoenix traces, so we could compare cache reuse, estimated cost, and latency on the…
7 min readBenchmarks, agent evaluations, and experiments.
View all research
Decision model benchmark: Jev, Kev, Liquid d1, and more
On 15 September, TypeSafe released Jev, and within a few weeks about 10 companies had shipped something that works the same way. OpenAI, Liquid, Cloudflare, AWS and PostHog…
October 7, 2026 13 min read
Prompt caching benchmark: high cache reuse doesn’t always mean lower cost
We benchmarked prompt caching across DeepSeek, GLM, GPT, and Claude using Harbor evals and Phoenix traces, so we could compare cache reuse, estimated cost, and latency on the…
October 2, 2026 7 min read
Are agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment.
September 29, 2026 9 min readDemos, workshops, and conference talks.
See all on YouTube
The Model Knew What to Say. The Harness Failed. | Laurie Voss
Laurie Voss (Head of Developer Relations at Arize AI) breaks down why the agent harness, specifically code-level guardrails, is the most critical boundary layer in modern software engineering.
InsightsGet the latest on AI & Observability
Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:
Real teams, shipping AI.
See allHow Booking.com scales AI observability with Arize
How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and…
Read moreHow LG Uplus is building better AI customer service agents with evaluation-driven development
How LG Uplus uses Arize AX to build evaluation-driven AI contact center agents — combining production traces, user feedback, and domain expertise to continuously improve customer service for…
Read moreHow Tripadvisor is building the AI product development lifecycle for agentic travel
Tripadvisor VP of Data and AI Rahul Todkar on building a production AI lifecycle for traditional ML and agentic travel products with unified observability, traces, evals, and governance…
Read moreCatch up on everything you missed.
See all-
Agent EngineeringIntroducing Arize AX MCP: When to use MCP, CLI, or skills
MCP, CLI, and skills are three ways to extend an agent with the same platform. The question is no longer which one wins. It is where the agent… Fuad Ali Dirk Brand Sragvi Vadali October 2, 2026 11 min read -
Agent EvaluationClaude’s hillclimb loop for AI agents: start with production traces
Anthropic's Claude hillclimb loop designs evals and improves agents one change at a time. Here's how to run that loop against production traces in Arize AX. Jim Bennett October 2, 2026 13 min read -
Agent EngineeringHow we built long-term memory for Alyx: why we chose a file over a knowledge graph
How we built long-term memory for Alyx: why we chose one 8,000-character file over retrieval and knowledge graphs, and how we tested it. Priyan Jindal Dheeraj Bandaru October 1, 2026 27 min read -
Agent EngineeringAlyx now remembers your work across sessions with long-term memory
Alyx now remembers project goals, conventions, and decisions across sessions in Arize AX, helping you continue work on traces, evals, and experiments. Chris Cooning Dheeraj Bandaru October 1, 2026 3 min read -
Agent EvaluationNVIDIA proposes an AI agent kill switch in silicon after a year of sandbox escapes
This year agents at OpenAI, Anthropic, and Google escaped environments meant to contain them. What decided severity was time-to-detect and time-to-kill, from 12 minutes to seven months. NVIDIA's… Jim Bennett September 30, 2026 13 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.