-
AI EvaluationOpenAI’s Santosh Vempala Explains Why Language Models Hallucinate
In our latest AI research paper reading, we hosted Santosh Vempala, Professor at Georgia Tech and co-author of OpenAI’s paper, “Why Language Models Hallucinate.” This paper offers one… Julian Reeves October 24, 2025 4 min read -
AI EvaluationWhat Are the Top LLM Evaluation Tools?
AI agents and real-world applications of generative AI are debuting at an incredible clip this year, narrowing the time from AI research paper to industry application and propelling… David Burch October 23, 2025 2 min read -
AI EvaluationShould I Use the Same LLM for My Eval as My Agent? Testing Self-Evaluation Bias
Thanks to Aparna Dhinakaran and Elizabeth Hutton for their contributions to this piece. When building and testing AI agents, one practical question that arises is whether to use… Sanjana Yeddula October 8, 2025 10 min read -
AI EvaluationNew In Arize AX: Session and Trace Evals, Alyx’s Synthetic Data Generation, and more
September was a busy month product-wise for Arize AX, with updates to make AI agent engineering faster and easier. From session and trace evals to Alyx’s new synthetic… Sanjana Yeddula October 6, 2025 3 min read -
AI EvaluationTesting Binary vs Score Evals on the Latest Models
Thanks to Hamel Husain and Eugene Yan for reviewing this piece Evals are becoming the predominant approach for how AI engineers systematically evaluate the quality of the LLM… Aparna Dhinakaran Sri Chavali September 24, 2025 10 min read -
AI EvaluationAtropos Health’s Arjun Mukerji, PhD, Explains RWESummary: A Framework and Test for Choosing LLMs to Summarize Real-World Evidence (RWE) Studies
Large language models are increasingly used to turn complex study output into plain-English summaries. But how do we know which models are safest and most reliable for healthcare? … Jason Lopatecki September 19, 2025 2 min read -
AI EvaluationRise of the Agent Engineer: Trunk Tools’ Bobby Vinson
Trunk Tools is building the brain behind construction, transforming the $13 trillion construction industry. As a premier AI agent platform for the built environment, Trunk Tools deploys solutions… David Burch September 19, 2025 4 min read -
AI EvaluationBuilding a Multilingual Cypher Query Evaluation Pipeline
How to evaluate LLM performance across languages for complex cypher query generation using open source tools As organizations expand globally, the need for multilingual AI systems becomes critical.… Mohit Talniya September 9, 2025 12 min read -
AI EvaluationAI Evals Maven Course Homework: the Recipe Bot Workflow
AI Evals for Engineers & PMs is a popular, hands‑on Maven course led by Hamel Husain and Shreya Shankar. The course’s goal is simple: “teach a systematic workflow… Sri Chavali September 3, 2025 9 min read -
AI EvaluationAnnotation for Strong AI Evaluation Pipelines
This post walks through how human annotations fit into your evaluation pipeline in Phoenix, why they matter, and how you can combine them with evaluations to build a… Sanjana Yeddula August 21, 2025 4 min read -
AI EvaluationHow Handshake Deployed and Scaled 15+ LLM Use Cases In Under Six Months — With Evals From Day One
Handshake is the largest early-career network, specializing in connecting students and new grads with employers and career centers. It’s also an engineering powerhouse and innovator in applying AI… Aparna Dhinakaran Kyle Gallatin August 21, 2025 4 min read -
AI EvaluationSession-Level Evaluations with Arize AX
When evaluating AI applications, we often look at things like tool calls, parameters, or individual model responses. While this span-level evaluation is useful, it doesn’t always capture the… Sanjana Yeddula August 19, 2025 3 min read -
AI EvaluationLLM Observability for AI Agents and Applications
The era of single-turn LLM calls is behind us. Today’s AI products are powered by increasingly autonomous agents — multi-step systems that plan, reason, use tools, and adapt… Sanjana Yeddula July 18, 2025 8 min read -
AI EvaluationSelf-Adapting Language Models: Paper Authors Discuss Implications
In a recent live AI research paper reading, the authors of the new paper Self-Adapting Language Models (SEAL) shared a behind-the-scenes look at their work, motivations, results, and… Jason Lopatecki July 8, 2025 4 min read -
AI EvaluationArize Observe 2025 – Product Releases
Arize Observe 2025 brought a wealth of new product releases, including a redesigned copilot, agent eval options, and state-of-the-art prompt optimization techniques. Check them all out below! Copilot… John Gilhuly June 25, 2025 7 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.