-
Agent EvaluationClaude’s hillclimb loop for AI agents: start with production traces
Anthropic's Claude hillclimb loop designs evals and improves agents one change at a time. Here's how to run that loop against production traces in Arize AX. Jim Bennett October 2, 2026 13 min read -
Agent EvaluationHow we built long-term memory for Alyx: why we chose a file over a knowledge graph
How we built long-term memory for Alyx: why we chose one 8,000-character file over retrieval and knowledge graphs, and how we tested it. Priyan Jindal Dheeraj Bandaru October 1, 2026 27 min read -
Agent EvaluationNVIDIA proposes an AI agent kill switch in silicon after a year of sandbox escapes
This year agents at OpenAI, Anthropic, and Google escaped environments meant to contain them. What decided severity was time-to-detect and time-to-kill, from 12 minutes to seven months. NVIDIA's… Jim Bennett September 30, 2026 13 min read -
Agent EvaluationWhat everyone was talking about at WeAreDevelopers World Congress North America
After 237 talks across eight stages, code review emerged as the bottleneck for AI-generated code. Here are the production eval, sandbox, context, and agent experience themes that kept… Laurie Voss September 29, 2026 16 min read -
Agent EvaluationAre agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment. Laurie Voss September 29, 2026 9 min read -
Agent EvaluationEvaluate production traces with Jev-as-a-Judge directly in Arize AX
Arize AX now integrates directly with TypeSafe AI, bringing native Jev-as-a-Judge evaluations into the Arize evaluation workflow. Chris Cooning September 28, 2026 3 min read -
Agent EvaluationWhat Jev’s probabilities reveal that repeated LLM judgments miss
I tested Jev and five LLM judges across ten Arize Phoenix evaluators to measure how often their answers changed alongside accuracy, cost, and latency. Elizabeth Hutton September 28, 2026 11 min read -
Agent EvaluationHow to build a remote evaluator in Arize AX with Jev
Evaluate agent responses with Jev and Arize AX. Build a FastAPI remote evaluator, return labels and scores, and choose a threshold for your data. Sally-Ann DeLucia Jim Bennett September 24, 2026 12 min read -
Agent EvaluationJev vs. LLM-as-a-Judge: Accuracy and cost benchmarks
We benchmarked Jev against Claude Opus 5 and GPT-5.6 Terra on accuracy, cost, and latency. Learn how threshold tuning changes hallucination detection. Laurie Voss September 23, 2026 11 min read -
Agent EvaluationReal-time LLM guardrails with Jev: comparing latency and cost
Compare Jev and GPT-5.4 nano for real-time LLM guardrails, with demo results on latency, cost, and checks on agent inputs, replies, and tool calls. Jim Bennett September 23, 2026 13 min read -
Agent EvaluationWhat changes when AI agents use your software
Daytona cofounder Ivan Burazin wants agents that can finish the job within the authority they have been given. His interview offers a starting point for examining how agents… Aaron Winston September 23, 2026 9 min read -
Agent EvaluationArize AX in September 2026: first-class sessions, Agent-as-a-Judge, and vision evals
Sessions are now a first-class unit of work in Arize AX. Annotate a whole conversation, queue it for review, and ask Alyx to filter for it — plus… Fuad Ali September 22, 2026 5 min read -
Agent EvaluationTypeSafe’s Jev: Can decision models replace LLM judges?
TypeSafe’s Jev classifies, scores, and routes without generating text — up to hundreds of times cheaper than an LLM judge. What that changes for evals, confidence routing, and… Laurie Voss September 18, 2026 10 min read -
Agent EvaluationHow to find and debug agent failures your evals are missing
Evals measure failures you know how to name. Arize Signal continuously reviews production traces to find recurring trajectory failures you do not, then turns them into evidence for… Aaron Winston September 17, 2026 13 min read -
Agent EvaluationAI agent regression testing with Agent Experiments in Arize AX
A cancellation-policy fix raised action safety and dropped average task completion from 0.89 to 0.72. This walkthrough shows how to regression-test agent changes with Agent Experiments in Arize… Nancy Chauhan Fuad Ali September 14, 2026 9 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.