Podcasts.
Accurate KV Cache Quantization with Outlier Tokens Tracing
Deploying large language models (LLMs) at scale is expensive—especially during inference. One of the biggest memory and performance…
Read the post
Scalable Chain of Thoughts via Elastic Reasoning
This paper introduces Elastic Reasoning, a novel framework designed to enhance the efficiency and scalability of large reasoning…
Read the post
AI Benchmark Deep Dive: Gemini 2.5 and Humanity’s Last Exam
Our latest paper reading provided a comprehensive overview of modern AI benchmarks, taking a close look at Google’s…
Read the post
Merge, Ensemble, and Cooperate! A Survey on Collaborative LLM Strategies
LLMs have revolutionized natural language processing, showcasing remarkable versatility and capabilities. But individual LLMs often exhibit distinct strengths…
Read the post
Introduction to OpenAI’s Realtime API
We break down OpenAI’s realtime API. Sally-Ann DeLucia and Aparna Dhinakaran cover how to seamlessly integrate powerful language…
Read the post
Swarm: OpenAI’s Experimental Approach to Multi-Agent Systems
As multi-agent systems grow in importance for fields ranging from customer support to autonomous decision-making, OpenAI has introduced…
Read the post
Google’s NotebookLM and the Future of AI-Generated Audio
In this paper read, Aman Khan and Harrison Chu explore NotebookLM’s unique features, including its ability to generate…
Read the post
Exploring OpenAI’s o1-preview and o1-mini
OpenAI recently released its o1-preview, which they claim outperforms GPT-4o on a number of benchmarks. These models are…
Read the post
Breaking Down Reflection Tuning: Enhancing LLM Performance with Self-Learning
A recent announcement on X boasted a tuned model with pretty outstanding performance, and claimed these results were…
Read the post
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Introduction This week’s paper presents a comprehensive study of the performance of various LLMs acting as judges. The…
Read the post
Breaking Down Meta’s Llama 3 Herd of Models
Introduction Meta just released Llama 3.1 405B–and according to them, it’s “the first openly available model that rivals…
Read the post
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
Introduction Chaining language model (LM) calls as composable modules is fueling a new way of programming, but ensuring…
Read the postDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.