Podcasts — page 2.
RAFT: Adapting Language Model to Domain Specific RAG
Introduction Where adapting LLMs to specialized domains is essential (e.g., recent news, enterprise private documents), we discuss a…
Read the post
LLM Interpretability and Sparse Autoencoders: Research from OpenAI and Anthropic
Introduction It’s been an exciting couple weeks for GenAI! Join us as we discuss the latest research from…
Read the post
Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment
Introduction We break down a paper, Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment.…
Read the post
Breaking Down EvalGen: Who Validates the Validators?
Introduction Due to the cumbersome nature of human evaluation and limitations of code-based evaluation, Large Language Models (LLMs)…
Read the post
Anthropic Claude 3
Introduction In this week’s Arize Community Paper Reading we dive into the latest buzz in the AI world—the…
Read the post
Reinforcement Learning in the Era of LLMs
Introduction This week, we explore Reinforcement Learning in the Era of LLMs: What is Essential? What is needed?…
Read the post
Sora: OpenAI’s Text-to-Video Generation Model
Introduction This week, we talk about the implications of Text-to-Video Generation and speculate as to the possibilities (and…
Read the post
Sora: OpenAI’s Text-to-Video Generation Model
Introduction This week, we discuss the implications of Text-to-Video Generation and speculate as to the possibilities (and limitations)…
Read the post
RAG vs Fine-Tuning
Introduction This week we discussed “RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture.” This paper…
Read the post
Phi-2 Model
Introduction With only 2.7 billion parameters, Phi-2 surpasses the performance of Mistral and Llama-2 models at 7B and…
Read the post
Mistral AI (Mixtral-8x7B): Performance, Benchmarks
Introduction For the last paper read of the year, Arize CPO & Co-Founder, Aparna Dhinakaran, is joined by…
Read the post
How to Prompt LLMs for Text-to-SQL
Introduction For this paper read, we’re joined by Shuaichen Chang, now an Applied Scientist at AWS AI Lab…
Read the postDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.