-
Agent EngineeringAgent harness vs. agent framework: why harnesses are replacing frameworks
Agent harnesses are replacing frameworks as the real product surface for reliable AI agents, shifting the work from prompt tuning to loops, tools, traces, evals, and operational metrics. Laurie Voss June 18, 2026 8 min read -
Agent EngineeringTwo labs started dreaming, and they built two different architectures
Anthropic and OpenAI both shipped 'dreaming' for AI memory in May and June 2026, and they built opposite architectures. A look at what each lab shipped, what the… Jim Bennett June 17, 2026 11 min read -
Agent EngineeringMemory is still a missing primitive: Cataloguing what the field is actually shipping
This week the field shipped four kinds of memory, and Apple paid Google a billion dollars a year for one of them. None of the four is what… Jim Bennett June 12, 2026 14 min read -
Agent EngineeringPostgresFS vs. SQL skills: should AI agents fake a filesystem?
Can an AI agent use a database as if it were a filesystem? Arize compared a Postgres-backed filesystem abstraction with a SQL skill and found that locality, accuracy,… Aparna Dhinakaran Sufjan Fana June 11, 2026 12 min read -
Agent EngineeringHow to detect credential theft in AI agent harness traces
In May 2026, a malicious version of a popular VS Code extension spent 18 minutes in the marketplace before anyone caught it. In that time it ran on… Nancy Chauhan June 9, 2026 14 min read -
Agent EngineeringBuilding the AI factory for self-improving agents: What’s new in Arize AX
Arize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. Jason Lopatecki Aparna Dhinakaran June 4, 2026 8 min read -
Agent EngineeringMicrosoft’s open trust stack runs on OpenInference
Microsoft's open trust stack for AI agents puts ASSERT and Agent Control Specification on top of OpenInference, connecting evaluation, runtime controls, and observability through a shared trace contract. Jim Bennett June 3, 2026 6 min read -
Agent EngineeringHow Hermes implements an open source agent harness architecture
Hermes from NousResearch is a strong open-source agent harness. This post examines how its runtime loop, context management, tool scoping, session infrastructure, and orchestration patterns map to a… Aparna Dhinakaran June 1, 2026 7 min read -
Agent EngineeringHow to build a better agent harness with traces and evals
Agents are easy to prototype and hard to improve. A repeatable loop of traces, evals, failed-span inspection, and targeted harness changes makes agent behavior easier to debug and… Aaron Winston May 29, 2026 14 min read -
Agent EngineeringWhat we learned testing 7 models under the same agent harness
Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on… Nancy Chauhan May 20, 2026 10 min read -
Agent EngineeringBuilding a self-improving agent on a context graph of human disagreement
You can build a measurably better agent from data you already have, without retraining a thing. The data is what your experienced humans do when they correct the… Jim Bennett May 19, 2026 12 min read -
Agent EngineeringCoding agent tracing and evaluation: An open source tool to improve AI coding workflows
Announcing coding harness tracing for observing, evaluating, and improving coding agent workflows across Claude Code, Cursor, Codex, GitHub Copilot, and Gemini CLI. Duncan McKinnon Chris Cooning Fuad Ali May 18, 2026 5 min read -
Agent EngineeringHow we use Alyx to build Alyx: How to build an AI agent feedback loop
How Arize uses Alyx to debug Alyx: searching dense traces, aggregating failures, triaging dogfooding issues, and closing the AI engineering feedback loop. Chris Cooning Sally-Ann DeLucia Priyan Jindal Jack Zhou May 13, 2026 10 min read -
Agent EngineeringAgent harnesses have an expiration date
A benchmark-driven look at why agent harnesses need adaptive finish logic as model behavior changes across Claude, GPT-4o, and Gemma. RL Nabors May 7, 2026 11 min read -
Agent EngineeringSwarm management in agent harnesses: owning long-running agents
As we have built our own harness management tools internally at Arize, and watched external systems like Devin @cognition start managing other Devins, managed agents at @AnthropicAI and… Aparna Dhinakaran May 4, 2026 11 min read
Don’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.