The Evaluator
Your go-to blog for insights on AI observability and evaluation.
Showing 251–260 of 458 posts (page 26 of 46)
Introducing Arize Copilot
If you used Microsoft Office in the early days, you probably remember Clippy. Clippy was an animated paper clip and go-to assistant for all things Microsoft Office. It provided users with tips, help, and shortcuts for using Microsoft Office applications. Well, it got us thinking…what if there was a Clippy but AI…for AI? Introducing Arize…
LlamaIndex’s Newly-Released Instrumentation Module + Phoenix Integration
Due to the black box nature of LLMs and the importance of tasks they’re being trusted to handle, intelligent monitoring and optimization tools are essential to ensure they operate efficiently and effectively. The integration of Arize Phoenix with LlamaIndex’s newly released instrumentation module offers developers unprecedented power to fine-tune performance, diagnose issues, and enhance the…
RAFT: Adapting Language Model to Domain Specific RAG
Introduction Where adapting LLMs to specialized domains is essential (e.g., recent news, enterprise private documents), we discuss a paper that asks how we adapt pre-trained LLMs for RAG in specialized domains. SallyAnn DeLucia is joined by Sai Kolasani, researcher at UC Berkeley’s RISE Lab (and Arize AI Intern), to talk about his work on RAFT:…
Sign up for our newsletter, The Evaluator — and stay in the know with updates and new resources:
Managing and Monitoring Your Open Source LLM Applications
LLMs are all the rage at the moment, and the APIs of closed source models like GPT-4 have made it easier than ever to leverage the power of AI. However, for a lot of regulated industries these closed source models are not an option. Luckily there is a plethora of open source alternatives, like Llama…
LLM Interpretability and Sparse Autoencoders: Research from OpenAI and Anthropic
Introduction It’s been an exciting couple weeks for GenAI! Join us as we discuss the latest research from OpenAI and Anthropic. We’re excited to chat about this significant step forward in understanding how LLMs work and the implications it has for deeper understanding of the neural activity of language models. We take a closer look at…
LLM Summarization: Getting To Production
Recently, I attended a workshop hosted by Arize AI’s Jason Lapatecki and Dat Ngo on large language model summarization covering common challenges with the use case and how to evaluate generated summaries. Drawing from this session and additional research, this article dives into the concept of LLM summarization – why it is important, primary summarization…
Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment
Introduction We break down a paper, Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment. Ensuring alignment (aka: making models behave in accordance with human intentions) has become a critical task before deploying LLMs in real-world applications. However, a major challenge faced by practitioners is the lack of clear guidance on evaluating…
How GetYourGuide Powers Millions of Real-Time Rankings with Production AI
This piece is co-authored by: Martin Jewell, Senior MLOps Engineer at GetYour Guide; Greg Chase, Machine Learning Solutions Engineer at Arize AI; and Mihir Mathur, Product Manager at Tecton GetYourGuide, a leading global online platform for discovering and booking travel experiences worldwide, makes over 30 million ranking predictions daily, shaping the journeys of millions of…
Arize AI Brings LLM Evaluation, Observability To Microsoft Azure AI Model Catalog
Generative AI is reshaping the modern enterprise. According to a recent survey, over half (61%) of developers say they plan to deploy LLM applications into production in the next 12 months or “as soon as possible.” However, challenges remain in getting a generative application from toy to production – and staying there. This week at…
Using Generative AI to Evaluate Bias in Speeches
Kansas City Chiefs kicker Harrison Butker recently sparked debate after delivering a commencement address to the 2024 graduating class at Benedictine College that touched on topics like gender roles, Pride Month, and President Joe Biden. With Butker’s words igniting intense debate and many offering their views across social media, we decided to investigate how AI…