Advanced LLM Evals: Creating an Eval from Scratch – Lessons from the Trenches with Bazaarvoice

Lou Kratz is Principal Research Engineer at Bazaarvoice, a service that is trusted by thousands of the world’s leading brands and retailers to drive revenue, extend reach, gain actionable insights, and create loyal advocates. Jason Lopatecki is CEO and Co-Founder of Arize AI In this session, Bazaarvoice will speak to their LLM use cases and share best practices around build LLM evals from scratch in partnership with Arize. This video is part two in a series on unpacking advanced LLM evaluation techniques and best practices formulated through rigorous testing — spanning retrieval, summarization, and hallucination — to help ensure production readiness. A must-attend for AI & ML engineers and data scientists.
Recommended resources
Prompt Playground
A prompt playground offers a UI to experiment with prompt templates, input variables, LLM models and LLM parameters.…
Read moreGolden Dataset: Role In Custom LLM Evals
Phoenix Community Challenge: Agents
November 1st – 30th Virtual We recently wrapped up our 6-week bootcamp on how to build AI…
Read moreDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.