The Evaluator
Your go-to blog for insights on AI observability and evaluation.
Showing 231–240 of 458 posts (page 24 of 46)
Creating and Validating Synthetic Datasets for LLM Evaluation & Experimentation
Thanks to John Gilhuly for his contributions to this piece. Looking for more on generating synthetic data and data evaluation? Book time with an Arize team member to discuss! Synthetic datasets are artificially created datasets that are designed to mimic real-world information. Unlike naturally occurring data, which is gathered from actual events or interactions, synthetic…
Evaluating an Image Classifier
Phoenix supports multi-modal evaluation and tracing. In this tutorial, we’ll take advantage of that to walk through the process of setting up an image classification experiment using Phoenix. This involves uploading a dataset, creating an experiment to classify the images, and evaluating the model’s accuracy. We’ll be using OpenAI’s GPT-4o-mini model for the classification task….
State of AI Engineering: Survey
Industries are racing to integrate large language models (LLMs) into their core operations. From better summarizing medical research to navigating complex case law, many early movers are seeing outsized benefits. The professionals building LLM systems — AI engineers, developers, data scientists, and industry leaders — play a pivotal role in shaping this future. This survey…
Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:
How To Set Up CrewAI Observability
Why Observability is Important with CrewAI In the world of autonomous AI agents, the ability to monitor and evaluate their performance is the key to unlocking their full potential. Observability allows you to see exactly what your agents are doing, how they are performing, and where they can be improved. Just as visibility into KPIs…
Arize Release Notes: Aug 23, 2024
Welcome to our regular update on new releases, enhancements, and changes. What’s New Create Spaces Programmatically Users can now create spaces programmatically with graphQL. Online Evals Update We added support for three new LLM integrations for online tasks: Azure OpenAI, Bedrock, and Vertex / Gemini. Event-Based Snowflake Jobs Users can now trigger Snowflake queries using…
How Bazaarvoice Navigated the Challenges of Deploying an LLM App
Bazaarvoice, a top platform for user-generated content and social commerce, has leveraged AI for much of its history — and now has a pioneering LLM app in production. At Arize:Observe this year, we caught up with Lou Kratz, Principal Research Engineer at Bazaarvoice, to talk about how their team successfully worked through the challenges that…
Trace Your Haystack Application
Haystack is an open-source framework for building LLM applications, retrieval-augmented generative pipelines and search systems that work intelligently over large document collections. Haystack makes it very easy to get off and running quickly with search and RAG LLM apps. Combining Haystack with Phoenix gives you the tools to not only build these apps quickly, but…
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Introduction This week’s paper presents a comprehensive study of the performance of various LLMs acting as judges. The researchers leverage TriviaQA as a benchmark for assessing objective knowledge reasoning of LLMs and evaluate them alongside human annotations which they find to have a high inter-annotator agreement. The study includes nine judge models and nine exam-taker…
How To Use Annotations To Collect Human Feedback On Your LLM Application
Liking Phoenix? Please consider giving us a ⭐ on Github! Cast your mind back to the early days of mainstream AI development – a whopping seven years ago. NVIDIA stock was just over $1 per share. “Transformers” meant Optimus Prime, not a word-changing technical leap forward. And the primary way of evaluating AI applications was…
How Atropos Health Accelerates Research with LLM Observability
Atropos Health aims to close the evidence gap to make it easier for physicians to have access to on-demand observational studies whenever needed. We caught up with Rebecca Hyde, Principal Data Scientist at Atropos Health, after her session at Arize:Observe in San Francisco. Hyde has over ten years of expertise in data science, public health,…