The Evaluator

Your go-to blog for insights on AI observability and evaluation.

Showing 91–100 of 458 posts (page 10 of 46)

Mastering Production RAG with Google ADK and Arize AX for Enterprise Knowledge Systems
Agent Engineering Agent Evaluation Agent Observability

Mastering Production RAG with Google ADK and Arize AX for Enterprise Knowledge Systems

Introduction Retrieval Augmented Generation (RAG) has become the cornerstone of enterprise AI, yet most organizations struggle with a critical challenge: building RAG systems that work reliably in production. While the promise is compelling, combining LLM reasoning with proprietary knowledge, the reality involves complex orchestration, sophisticated evaluation, and continuous monitoring that traditional frameworks don’t address. Google’s…

How America First Credit Union Built a GenAI “Decision Explainer” — With Tracing That Scales
AI Observability AI Product Quality Case Studies

How America First Credit Union Built a GenAI “Decision Explainer” — With Tracing That Scales

America First Credit Union is one of America’s largest independent credit unions, with 1.5 million members and more than $20 billion worth of deposits. As America First Credit Union scaled AI across its business, a new challenge emerged: business stakeholders still needed to explain decisions, but the “why” increasingly lived inside complex statistical models. To…

Closing the Loop: Coding Agents, Telemetry, and the Path to Self-Improving Software
Agent Engineering Agent Evaluation Agents

Closing the Loop: Coding Agents, Telemetry, and the Path to Self-Improving Software

2025 marked the widespread adoption of coding agents — harnesses that autonomously write, test, and debug changes to software with minimal human intervention. Products like Claude Code, Codex, Cursor, and Open Code are churning out unfathomable lines of code a day. A recent large-scale study of GitHub repositories estimated that 16 to 23 percent of…

Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:

Inside Typeform’s AI Agent Stack
Agent Engineering Agent Evaluation AI Evaluation

Inside Typeform’s AI Agent Stack

Typeform is building generative AI experiences to help customers create better forms faster and to make collecting insights feel more natural and useful end-to-end. In this Q&A, Marta Lorens, Senior Data Scientist at Typeform, shares how Typeform thinks about agentic and gen-AI use cases, why evaluation is a core part of the product experience, and…

CUGA Agent: From Benchmarks to Business Impact of IBM’s Generalist Agent
Agent Engineering Agent Evaluation AI Engineering

CUGA Agent: From Benchmarks to Business Impact of IBM’s Generalist Agent

This paper reading features several of the researchers — including Segev Shlomov‏ (PhD), Ido Levy, Asaf Adi, and Avi Yaeli — behind the widely acclaimed paper “From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production.” The paper reports IBM’s experience developing and piloting the Computer Using Generalist Agent (CUGA), which has been…

Top Generative AI Conferences In 2026 for Engineers
AI Evaluation Security & Governance

Top Generative AI Conferences In 2026 for Engineers

GenAI stacks are shifting fast enough that staying current is an ongoing project, not a quarterly refresh. The hard part is separating durable engineering practices (evals, reliability, cost controls, security) from transient tooling churn, so this list prioritizes events with repeatable signal for agent builders and AI engineers. Selection bias: conferences that help you ship…

New In Arize AX: January 2026 Updates
AI Evaluation LLM Evaluation Product Releases

New In Arize AX: January 2026 Updates

Arize AX pushed out a lot of new updates in January 2026. From improved evaluator hub to custom prompt release labels, here are some highlights. Evaluator Hub: Reusable Evaluators We’re excited to introduce the Evaluator Hub — a centralized place to create, version, and reuse evaluators across all your evaluation tasks. Why reusable evaluators? Previously,…

How Nebulock Democratizes Threat Hunting
Case Studies Security & Governance

How Nebulock Democratizes Threat Hunting

Nebulock is on a mission to democratize threat hunting. Instead of relying only on deterministic rules or reacting to alerts as they come in, the team builds AI agents that actively hunt inside an organization’s environment and surface pernicious threats with clear, actionable outcomes. These agents can run automatically — triggered by fresh threat intelligence…

Why AI Agents Break: A Field Analysis of Production Failures
Agent Observability Agents AI Observability

Why AI Agents Break: A Field Analysis of Production Failures

As AI agents enter production environments, they face conditions their training does not cover. These systems generate fluent output, yet operational work demands exact action. Small ambiguities compound fast when agents interact with live data and user constraints. An agent hallucinates a parameter because a field name appears valid. It fabricates a refund policy to…

OWASP Top 10 for Agentic Applications: Compliance Guide
Agent Engineering Agent Observability AI Observability

OWASP Top 10 for Agentic Applications: Compliance Guide

This guide maps the OWASP Agentic Security Initiative (ASI) top ten risks to specific Arize AX observability features and metrics you should implement to detect, monitor, and mitigate threats in your agentic AI systems. The OWASAP Agentic Security Initiative is a specialized project under the Open Web Application Security Project (OWASP) Generative AI Security Project….