Blog — page 11.
LLM-as-a-Judge: Example of How To Build a Custom Evaluator Using a Benchmark Dataset
When To Build Custom Evaluators Arize-Phoenix ships with pre-built evaluators that are tested against benchmark datasets and tuned…
Read the post
adb Database: Realtime Ingestion At Scale
We put out our first blog on the introducing the Arize database – adb – in the beginning…
Read the post
New In Arize AX: Prompt Learning, Arize Tracing Assistant, and Multiagent Visualization
July was a big month for Arize AX, with updates to make AI and agent engineering much easier.…
Read the post
A Watermark for Large Language Models
In our latest live AI research papers community reading, the primary author of the popular paper A Watermark…
Read the post
Unlocking Safer AI: Your Two-Part Field Guide
Large language models are reshaping how we build products — and how adversaries try to break them. To…
Read the post
LLM Observability for AI Agents and Applications
The era of single-turn LLM calls is behind us. Today’s AI products are powered by increasingly autonomous agents…
Read the post
Prompt Learning: Using English Feedback to Optimize LLM Systems
Applications of reinforcement learning (RL) in AI model building has been a growing topic over the past few…
Read the post
Self-Adapting Language Models: Paper Authors Discuss Implications
In a recent live AI research paper reading, the authors of the new paper Self-Adapting Language Models (SEAL)…
Read the post
Meet Alyx: Arize’s Evolving AI Agent
We’re excited to introduce Alyx, the next evolution in Arize’s intelligent assistant. You might remember our first iteration…
Read the post
Introducing adb: Arize’s Proprietary OLAP Database
Earlier this month, we rolled out real‑time ingestion support to every Arize AX workspace—paid and free. With that…
Read the post
The Illusion of Thinking: What the Apple AI Paper Says About LLM Reasoning
A recent paper from Apple researchers—The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via…
Read the post
Accurate KV Cache Quantization with Outlier Tokens Tracing
Deploying large language models (LLMs) at scale is expensive—especially during inference. One of the biggest memory and performance…
Read the postDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.