This paper reading features several of the researchers — including Segev Shlomov (PhD), Ido Levy, Asaf Adi, and Avi Yaeli — behind the widely acclaimed paper “From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production.” The paper reports IBM’s experience developing and piloting the Computer Using Generalist Agent (CUGA), which has been open-sourced for the community. CUGA adopts a hierarchical planner–executor architecture with strong analytical foundations, achieving state-of-the-art performance on AppWorld and WebArena. Beyond benchmarks, it was evaluated in a pilot within the Business-Process-Outsourcing talent acquisition domain, addressing enterprise requirements for scalability, auditability, safety, and governance.
CUGA Agent: From Benchmarks to Business Impact of IBM’s Generalist Agent
This paper reading features several of the researchers — including Segev Shlomov (PhD), Ido Levy, Asaf Adi, and Avi Yaeli — behind the widely acclaimed paper “From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production.” The paper reports IBM’s experience developing and piloting the Computer Using Generalist Agent (CUGA), which has been...
Get the latest on AI & Observability
Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:
Keep building the quality bar.
See allWhat everyone was talking about at WeAreDevelopers World Congress North America
After 237 talks across eight stages, code review emerged as the bottleneck for AI-generated code. Here are the production eval, sandbox, context, and agent experience themes that kept…
16 min readAre agent harnesses dying? What harness distillation changes
Harness distillation can train general scaffolding into a model. The harness that remains is the part tied to your tools, data, users, and environment.
9 min readAnthropic says it fixed Claude’s writing. I ran the evals to check.
Opus 5.5 dropped em dashes from 12.9 per 1,000 words to two in 57,000 words, and cut the rest of the Claudisms in half. Better, but not fixed.
12 min readDon’t ship vibes.
Arize gives AI teams observability and evals to understand and improve agent performance.