The Evaluator

Your go-to blog for insights on AI observability and evaluation.

Showing 131–140 of 458 posts (page 14 of 46)

New In Arize AX: Session and Trace Evals, Alyx’s Synthetic Data Generation, and more
Agent Engineering Agent Evaluation Agent Observability

New In Arize AX: Session and Trace Evals, Alyx’s Synthetic Data Generation, and more

September was a busy month product-wise for Arize AX, with updates to make AI agent engineering faster and easier. From session and trace evals to Alyx’s new synthetic data generation skill, there is a lot to catch up on and try out. Session and Trace Evals Session- and trace-level evaluations are available across all Arize…

Testing Binary vs Score Evals on the Latest Models
AI Evaluation LLM Evals LLM Evaluation

Testing Binary vs Score Evals on the Latest Models

Thanks to Hamel Husain and Eugene Yan for reviewing this piece Evals are becoming the predominant approach for how AI engineers systematically evaluate the quality of the LLM generated outputs. Despite this, teams have wildly different methods for how they are defining their evals – some use strictly boolean, while others use a variation of…

Atropos Health’s Arjun Mukerji, PhD, Explains RWESummary: A Framework and Test for Choosing LLMs to Summarize Real-World Evidence (RWE) Studies
AI Evaluation

Atropos Health’s Arjun Mukerji, PhD, Explains RWESummary: A Framework and Test for Choosing LLMs to Summarize Real-World Evidence (RWE) Studies

Large language models are increasingly used to turn complex study output into plain-English summaries. But how do we know which models are safest and most reliable for healthcare?  In this most recent community AI research paper reading, Arjun Mukerji, PhD – Staff Data Scientist at Atropos Health – walks us through RWESummary, a new benchmark…

Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:

Rise of the Agent Engineer: Trunk Tools’ Bobby Vinson
Agent Engineering Agent Evaluation Agents

Rise of the Agent Engineer: Trunk Tools’ Bobby Vinson

Trunk Tools is building the brain behind construction, transforming the $13 trillion construction industry. As a premier AI agent platform for the built environment, Trunk Tools deploys solutions that streamline construction data management, automate tedious and repetitive tasks, and minimize waste. In this interview, we catch up with Trunk Tools AI Evaluation Engineer Bobby Vinson…

adb Benchmarks
AI Observability

adb Benchmarks

In launching adb (Arize database) we wanted to benchmark adb both internally as a database and at the system level in our application. Our goal is to show both the performance of the database that powers the application and the end application delivered user experience.  The benchmark tests cover these areas: Dataset Upload Programmatic  Dataset…

Orchestrator-Worker Agents: A Practical Comparison of Common Agent Frameworks
Agent Engineering Agents AI Engineering

Orchestrator-Worker Agents: A Practical Comparison of Common Agent Frameworks

— Technical deep dive inspired by Anthropic’s “Building Effective Agents” In this piece, we’ll take a close look at the orchestrator–worker agent workflow. We’ll unpack its challenges and nuances, then compare how leading frameworks – Agno, Autogen, CrewAI, OpenAI, LangGraph, and Mastra – approach and implement this pattern. Orchestrator-Worker Architecture The orchestrator-worker architecture is designed…

Building a Multilingual Cypher Query Evaluation Pipeline
AI Evaluation LLM Evaluation Open Source

Building a Multilingual Cypher Query Evaluation Pipeline

How to evaluate LLM performance across languages for complex cypher query generation using open source tools As organizations expand globally, the need for multilingual AI systems becomes critical. But how do you evaluate whether your language model can handle business questions in non-English languages and still generate correct database queries? In this post, we’ll walk…

Verizon’s Stan Miasnikov Walks Through His Latest Paper On Inter-Agent Communication
Agent Engineering AI Engineering

Verizon’s Stan Miasnikov Walks Through His Latest Paper On Inter-Agent Communication

In a recent Arize community AI research paper reading, we had the honor to host Stan Miasnikov – Distinguished Engineer, AI/ML Architecture, Consumer Experience at Verizon – to highlight the findings of a series of papers including his most recent titled “Category-Theoretic Analysis of Inter-Agent Communication and Mutual Understanding Metric in Recursive Consciousness.” The paper…

New In Arize AX: Experiment Comparisons, Better Data Visualization, and a Dedicated Agent Graph Tab
Agent Engineering Agent Observability AI Engineering

New In Arize AX: Experiment Comparisons, Better Data Visualization, and a Dedicated Agent Graph Tab

August was a busy month, with lots of updates from the engineering team to make agent engineering easier. From previewing examples in the UI to a dedicated agent graph tab to richer charting for experiments, there is a lot to explore. Here are some highlights on what we shipped. Experiments & Data Visualization Improvements The…

NVIDIA’s Peter Belcak Distills Why Small Language Models are the Future of Agentic AI
Agent Engineering AI Engineering

NVIDIA’s Peter Belcak Distills Why Small Language Models are the Future of Agentic AI

In our most recent AI research paper community reading, we had the privilege of hosting Peter Belcak – an AI Researcher working on the reliability and efficiency of agentic systems at NVIDIA – who walked us through his new paper making the rounds in AI circles titled “Small Language Models are the Future of Agentic…