Synthetic Data Generation

Amber Roberts is a data scientist and machine learning engineer at Arize AI and leads the company's learning and development efforts. This video is part four in a series on unpacking advanced LLM evaluation techniques and best practices formulated through rigorous testing — spanning retrieval, summarization, and hallucination — to help ensure production readiness. A must-attend for AI & ML engineers and data scientists. This session covers classes of LLM evaluations.
See the rest of the series and sign up for future events:
https://arize.com/advanced-llm-evaluation
Learn more about LLM evaluation:
https://arize.com/blog-course/llm-evaluation-the-definitive-guide/
Recommended resources
LLM Evaluation: Everything You Need To Run, Benchmark LLM Evals
This piece is co-authored by senior machine learning engineer Ilya Reznik Large language models (LLMs) are an incredible…
Start the course
The Definitive LLM Observability Checklist
According to a recent survey, only 30.1% of teams deploying LLMs have implemented observability despite large majorities wanting…
Read more
LLM Observability: One Small Step for Spans, One Giant Leap for Span-Kinds
Troubleshooting Generative AI Models in Production As AI teams navigate the best practices for implementing large language models…
Start the courseDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.