-
Agent EngineeringMemory is still a missing primitive: Cataloguing what the field is actually shipping
This week the field shipped four kinds of memory, and Apple paid Google a billion dollars a year for one of them. None of the four is what… Jim Bennett June 12, 2026 14 min read -
Agent EngineeringPostgresFS vs. SQL skills: should AI agents fake a filesystem?
Can an AI agent use a database as if it were a filesystem? Arize compared a Postgres-backed filesystem abstraction with a SQL skill and found that locality, accuracy,… Aparna Dhinakaran Sufjan Fana June 11, 2026 12 min read -
Agent EngineeringHow to detect credential theft in AI agent harness traces
In May 2026, a malicious version of a popular VS Code extension spent 18 minutes in the marketplace before anyone caught it. In that time it ran on… Nancy Chauhan June 9, 2026 14 min read -
Agent EngineeringBuilding the AI factory for self-improving agents: What’s new in Arize AX
Arize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. Jason Lopatecki Aparna Dhinakaran June 4, 2026 8 min read -
Agent EngineeringMicrosoft’s open trust stack runs on OpenInference
Microsoft's open trust stack for AI agents puts ASSERT and Agent Control Specification on top of OpenInference, connecting evaluation, runtime controls, and observability through a shared trace contract. Jim Bennett June 3, 2026 6 min read -
Agent EngineeringHow Hermes implements an open source agent harness architecture
Hermes from NousResearch is a strong open-source agent harness. This post examines how its runtime loop, context management, tool scoping, session infrastructure, and orchestration patterns map to a… Aparna Dhinakaran June 1, 2026 7 min read -
Agent EngineeringHow to build a better agent harness with traces and evals
Agents are easy to prototype and hard to improve. A repeatable loop of traces, evals, failed-span inspection, and targeted harness changes makes agent behavior easier to debug and… Aaron Winston May 29, 2026 14 min read -
Agent EngineeringWhat we learned testing 7 models under the same agent harness
Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on… Nancy Chauhan May 20, 2026 10 min read -
Agent EngineeringBuilding a self-improving agent on a context graph of human disagreement
You can build a measurably better agent from data you already have, without retraining a thing. The data is what your experienced humans do when they correct the… Jim Bennett May 19, 2026 12 min read -
Agent EngineeringCoding agent tracing and evaluation: An open source tool to improve AI coding workflows
Announcing coding harness tracing for observing, evaluating, and improving coding agent workflows across Claude Code, Cursor, Codex, GitHub Copilot, and Gemini CLI. Duncan McKinnon Chris Cooning Fuad Ali May 18, 2026 5 min read -
Agent EngineeringHow we use Alyx to build Alyx: How to build an AI agent feedback loop
How Arize uses Alyx to debug Alyx: searching dense traces, aggregating failures, triaging dogfooding issues, and closing the AI engineering feedback loop. Chris Cooning Sally-Ann DeLucia Priyan Jindal Jack Zhou May 13, 2026 10 min read -
Agent EngineeringAgent harnesses have an expiration date
A benchmark-driven look at why agent harnesses need adaptive finish logic as model behavior changes across Claude, GPT-4o, and Gemma. RL Nabors May 7, 2026 11 min read -
Agent EngineeringSwarm management in agent harnesses: owning long-running agents
As we have built our own harness management tools internally at Arize, and watched external systems like Devin @cognition start managing other Devins, managed agents at @AnthropicAI and… Aparna Dhinakaran May 4, 2026 11 min read -
Agent EngineeringMCP vs. CLI Skills for agents: what our eval found (and which you should use)
Twitter said pick a side. The eval said the question was wrong. Six months ago, MCP (model context protocol) was the hot new thing: tool usage with a… Laurie Voss May 1, 2026 10 min read -
Agent EngineeringUsing context graphs: build a data moat like Google’s using your enterprise data
Enterprise software is on the verge of its first compounding data loop, the same kind of self-reinforcing mechanism that built the most valuable consumer businesses of the last… Jim Bennett Laurie Voss April 29, 2026 8 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.