Resource Hub

Inside Cursor’s agent factory: how it verifies AI-written code

Inside Cursor’s agent factory: how it verifies AI-written code

As background agents take on more implementation work, Cursor is rebuilding the software development lifecycle around risk scores, developer-like environments, video evidence, and review systems that learn from every human correction.

Why do you need evals for AI agents?

Why do you need evals for AI agents?

Agent failures hide inside correct-looking outputs. Learn why evals are the only mechanism that catches them, and how to wire them into dev, CI, and production.

LLM-as-a-Judge: When should you use it?

LLM-as-a-Judge: When should you use it?

LLM judges fit a narrow window. Learn the specific conditions that qualify a task, three prerequisites, and when a regex or compiler is the better choice.

What is AI engineering?

What is AI engineering?

AI engineering takes a capable model as a given and builds a reliable product around it.

The hidden economics of LLM evaluation costs

The hidden economics of LLM evaluation costs

Learn what drives LLM evaluation costs across offline tests and production, from judge tokens and trace volume to agent workflows, routing, and release decisions.

Kiro CLI observability: trace and evaluate agent changes with Arize Skills
Blog

Kiro CLI observability: trace and evaluate agent changes with Arize Skills

Use Arize Skills with Kiro CLI to trace coding-agent changes, build datasets from failures, run experiments, and validate prompts before shipping.

From human-operated agent development to systematic agent improvement
Blog

From human-operated agent development to systematic agent improvement

At Observe 2026, Jason Lopatecki and Aparna Dhinakaran described the shift from human-operated agent development to systematic agent improvement—and what builders should change in their stacks first.

How to measure AI productivity: From LLM token costs to business value with Arize AX
Blog

How to measure AI productivity: From LLM token costs to business value with Arize AX

AI productivity is best measured by connecting AI usage to validated downstream outcomes. Tokens, prompts, and generated lines show activity, but they do not prove value. A better measurement model tracks the cost of AI work, scores the quality and task success of that work, then joins each trace to outcomes such as merged PRs,...

How do you make an LLM, anyway? Microsoft just published a textbook.

How do you make an LLM, anyway? Microsoft just published a textbook.

Microsoft published a 109-page technical report on MAI-Thinking-1. Here’s the abbreviated version of how a modern lab actually trains a frontier reasoning model — from scraping the web to reinforcement learning, judges, and anti-cheating.

3 production patterns for AI agents and how to evaluate each one

3 production patterns for AI agents and how to evaluate each one

A local coding agent, an in-app customer assistant, and an AI SRE triaging production logs may all use the same model class—but not the same harness, eval plan, or rollout risk. Mastra CEO Sam Bhagwat breaks down the three production patterns and how to evaluate each one.

What is a loop in AI engineering, anyway?
Blog

What is a loop in AI engineering, anyway?

The AI engineering world is using “loop” to describe several different agent architectures. This post maps execution loops, task loops, product loops, system loops, and the human oversight loop that controls them.

Trace before you migrate: Measuring Kubernetes bottlenecks in AI agent sandboxes

Trace before you migrate: Measuring Kubernetes bottlenecks in AI agent sandboxes

Kubernetes is strong for long-lived services, but it is often a poor default for short-lived agent sandboxes. Trace sandbox creation, tool execution, eval latency, and full trajectory time before you migrate runtimes.

No results found. Try a different filter or search term.