The Evaluator

Your go-to blog for insights on AI observability and evaluation.

Showing 171–180 of 458 posts (page 18 of 46)

40 Large Language Model Benchmarks and The Future of Model Evaluation
LLM Evals LLM Evaluation

40 Large Language Model Benchmarks and The Future of Model Evaluation

With the accelerated development of GenAI, there is a particular focus on its testing and evaluation, resulting in the release of several LLM benchmarks. Each of these benchmarks tests the LLM’s different capabilities–but are they sufficient for a complete real-world performance evaluation? This blog will discuss some of the most popular LLM benchmarks for evaluating…

Building and Deploying Observable AI Agents with Google Agent Framework and Arize
Agent Observability Agents AI Observability

Building and Deploying Observable AI Agents with Google Agent Framework and Arize

Co-authored by Ali Arsanjani, Director of Applied AI Engineering at Google Cloud 1. Introduction: The Dawn of the Agentic Era We have entered into a new era of AI innovation and adoption, driven by three transformative trends: Multimodality, Multi-agent Systems,and Agentic Workflows. These trends have converged to make AI agents a game-changing technology for organizations…

Embracing Google’s Agent-To-Agent (A2A) Protocol
Agent Engineering Agents AI Engineering

Embracing Google’s Agent-To-Agent (A2A) Protocol

We’re excited to announce that Arize AI is partnering with Google as a launch partner for the Agent Interop Protocol (A2A), an open standard enabling seamless communication between AI agents across different platforms and organizational boundaries. What is Agent Interop Protocol (A2A)? The Agent Interop Protocol (A2A) addresses a fundamental challenge in today’s AI ecosystem:…

Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:

Tracing and Evaluating Gemini Audio with Arize
AI Observability LLM Evals LLM Observability

Tracing and Evaluating Gemini Audio with Arize

Google’s Gemini models represent a powerful leap forward in multimodal AI, particularly in their ability to process and transcribe audio content with remarkable accuracy. However, even advanced models require robust monitoring and evaluation frameworks to ensure consistent quality in production environments. This is where Arize’s tracing and evaluation capabilities become invaluable. By combining Gemini’s audio…

AI Benchmark Deep Dive: Gemini 2.5 and Humanity’s Last Exam
AI Evaluation Research

AI Benchmark Deep Dive: Gemini 2.5 and Humanity’s Last Exam

Our latest paper reading provided a comprehensive overview of modern AI benchmarks, taking a close look at Google’s recent Gemini 2.5 release and its performance on key evaluations, notably the challenging Humanity’s Last Exam (HLE). For those who missed the live session, we’ve compiled the essential highlights and key takeaways. This recap covers Gemini 2.5’s…

Model Context Protocol (MCP) from Anthropic
Agent Engineering Agents AI Engineering

Model Context Protocol (MCP) from Anthropic

Want to learn more about Anthropic’s groundbreaking Model Context Protocol (MCP)? We break down how this open standard is revolutionizing AI by enabling seamless integration between LLMs and external data sources, fundamentally transforming them into capable, context-aware agents. We explore the key benefits of MCP, including enhanced context retention across interactions, improved interoperability for agentic…

Self-Improving Agents: Automating LLM Performance Optimization using Arize and NVIDIA NeMo
Agent Engineering Agents Integrations

Self-Improving Agents: Automating LLM Performance Optimization using Arize and NVIDIA NeMo

Enterprises face a critical challenge in keeping their LLM models accurate and reliable over time. Traditional model improvement approaches are slow, manual, and reactive, making it difficult to scale and adapt to evolving data patterns. The Arize integration of NVIDIA NeMo empowers AI teams with an automated, self-improving AI data flywheel to enhance LLM performance….

Prompt Optimization Techniques
Prompt Engineering

Prompt Optimization Techniques

LLMs are powerful tools, but their performance is heavily influenced by how prompts are structured. The difference between an effective and ineffective prompt can determine whether a model produces accurate responses or fails to generalize to new inputs. Prompt optimization is the process of refining prompts to improve model outputs. This can involve adjusting wording,…

Prompt Management from First Principles
AI Observability LLM Observability Open Source

Prompt Management from First Principles

How we built a holistic prompt management system that preserves developer freedom Unlike traditional software, where code execution follows predictable paths, LLM applications are inherently non-deterministic. Their behavior is shaped by natural language inputs—’prompts’—which can produce vastly different results with even minor adjustments. Superficial changes in wording or style can yield dramatically different results, making…

How We Scaled Support in Arize Copilot Without Slowing Down
AI Engineering AI Product Quality

How We Scaled Support in Arize Copilot Without Slowing Down

Arize Copilot has always had a clear vision: to empower AI engineers and data scientists to spend less time on repetitive tasks and more time building innovative applications. Copilot streamlines workflows, automates debugging, and surfaces actionable insights—all designed to help users move faster and achieve more. With Copilot, we’ve always prioritized putting its capabilities where…