Classes of LLM Evaluations: A Deep Dive

Aparna Dhinakaran is Co-Founder and Chief Product Officer of Arize AI; Dat Ngo is an ML Solutions Architect at Arize AI. This video is part three in a series on unpacking advanced LLM evaluation techniques and best practices formulated through rigorous testing — spanning retrieval, summarization, and hallucination — to help ensure production readiness. A must-attend for AI & ML engineers and data scientists. This session covers classes of LLM evaluations.
Recommended resources
The Definitive LLM Observability Checklist
According to a recent survey, only 30.1% of teams deploying LLMs have implemented observability despite large majorities wanting…
Read more
How Handshake Deployed and Scaled 15+ LLM Use Cases In Under Six Months — With Evals From Day One
Handshake is the largest early-career network, specializing in connecting students and new grads with employers and career centers.…
Read the post
40 Large Language Model Benchmarks and The Future of Model Evaluation
With the accelerated development of GenAI, there is a particular focus on its testing and evaluation, resulting in…
Read the postDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.