The Evaluator
Your go-to blog for insights on AI observability and evaluation.
Showing 271–280 of 458 posts (page 28 of 46)
What Does It Take To Pioneer Successful LLM Applications In Healthcare and the Life Sciences?
Peter Leimbigler is a Data Science Team Leader within the Consulting practice at Klick Health. As the largest independent commercialization partner in its industry, Klick pioneers new AI-powered applications for clients across life sciences, pharmaceuticals, medical device, and consumer health to accelerate growth and improve experiences and outcomes for patients and consumers. That’s especially true…
Evaluating and Analyzing Your RAG Pipeline with Ragas
This article is co-authored by Mikyo King, Founding Engineer and Head of Open Source at Arize AI, and Xander Song, AI Engineer at Arize AI Building a baseline for a RAG pipeline is not usually difficult, but enhancing it to make it suitable for production and ensuring the quality of your responses is almost always…
Evaluating the Generation Stage in RAG
In retrieval-augmented generation (RAG), retrieval often steals the spotlight, while the generation stage receives less attention. To address this gap, we conducted a series of tests to see how different models handle the generation phase, and the results surprised us: Anthropic’s Claude outperformed OpenAI’s GPT-4 in generating responses. This outcome was unexpected, as GPT-4 usually…
Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:
RAG vs Fine-Tuning
Introduction This week we discussed “RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture.” This paper that explores a pipeline for Fine-tuning and RAG, and presents the tradeoffs of both for multiple popular LLMs, including Llama 2-13B, GPT-3.5, and GPT-4. The authors propose a pipeline that consists of multiple stages, including extracting information…
Phi-2 Model
Introduction With only 2.7 billion parameters, Phi-2 surpasses the performance of Mistral and Llama-2 models at 7B and 13B parameters on various aggregated benchmarks. Notably, it achieves better performance compared to 25x larger Llama-2-70B model on multi-step reasoning tasks, i.e., coding and math. Furthermore, Phi-2 matches or outperforms the recently-announced Google Gemini Nano 2, despite…
Diving Into Enterprise Data Strategy With Samsung Research’s Prashanth Rajendran
On a recent earnings call, Microsoft CEO Satya Nadella observed: “Every AI app starts with data and having a comprehensive data and analytics platform is more important than ever.” While much has changed over the past year with the emergence of generative AI, data quality is still a fundamental building block of any enterprise –…
Top AI Conferences of 2024: Generative AI and Beyond
Psst…Looking for 2025 conferences? Click here. As we prepare for another remarkable year in artificial intelligence, the importance of keeping up with the latest advancements, trends, and breakthroughs cannot be overstated. In 2024, leaders across the global AI landscape will participate in conferences across the globe that will predict and help shape the future. Whether…
Evaluate RAG with LLM Evals and Benchmarking
Recently, I attended a workshop organized by Arize AI titled “RAG Time! Evaluate RAG with LLM Evals and Benchmarking.” Hosted by Amber Roberts – ML Growth Lead at Arize AI, and Mikyo King – Head of Open Source at Arize AI, the talks provided valuable insights into an important field of study. Miss the event?…
Mistral AI (Mixtral-8x7B): Performance, Benchmarks
Introduction For the last paper read of the year, Arize CPO & Co-Founder, Aparna Dhinakaran, is joined by a Dat Ngo (ML Solutions Architect) and Aman Khan (Group Product Manager) for an exploration of the new kids on the block: Gemini and Mixtral-8x7B. There’s a lot to cover, so this week’s paper read is Part…
Why Enterprise Executives Should Be Hip To LLMOps Tools Heading Into the New Year
From better customer service to more rapid drug discovery, generative AI is quickly reshaping industries. According to a recent survey, 61.7% of enterprise engineering teams now have or are planning to have a large language model application deployed into the real world within a year – with over one in ten (14.7%) already in production,…