All resources

Blog — page 11.

Blog

LLM-as-a-Judge: Example of How To Build a Custom Evaluator Using a Benchmark Dataset

When To Build Custom Evaluators Arize-Phoenix ships with pre-built evaluators that are tested against benchmark datasets and tuned…

Read the post
Blog

adb Database: Realtime Ingestion At Scale

We put out our first blog on the introducing the Arize database – adb – in the beginning…

Read the post
Blog

New In Arize AX: Prompt Learning, Arize Tracing Assistant, and Multiagent Visualization

July was a big month for Arize AX, with updates to make AI and agent engineering much easier.…

Read the post
Blog

A Watermark for Large Language Models

In our latest live AI research papers community reading, the primary author of the popular paper A Watermark…

Read the post
Blog

Unlocking Safer AI: Your Two-Part Field Guide

Large language models are reshaping how we build products — and how adversaries try to break them. To…

Read the post
Blog

LLM Observability for AI Agents and Applications

The era of single-turn LLM calls is behind us. Today’s AI products are powered by increasingly autonomous agents…

Read the post
Blog

Prompt Learning: Using English Feedback to Optimize LLM Systems

Applications of reinforcement learning (RL) in AI model building has been a growing topic over the past few…

Read the post
Blog

Self-Adapting Language Models: Paper Authors Discuss Implications

In a recent live AI research paper reading, the authors of the new paper Self-Adapting Language Models (SEAL)…

Read the post
Blog

Meet Alyx: Arize’s Evolving AI Agent

We’re excited to introduce Alyx, the next evolution in Arize’s intelligent assistant. You might remember our first iteration…

Read the post
Blog

Introducing adb: Arize’s Proprietary OLAP Database

Earlier this month, we rolled out real‑time ingestion support to every Arize AX workspace—paid and free. With that…

Read the post
Blog

The Illusion of Thinking: What the Apple AI Paper Says About LLM Reasoning

A recent paper from Apple researchers—The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via…

Read the post
Blog

Accurate KV Cache Quantization with Outlier Tokens Tracing

Deploying large language models (LLMs) at scale is expensive—especially during inference. One of the biggest memory and performance…

Read the post

Don’t ship vibes.

Arize gives AI teams observability and evals to understand and improve agent performance.