WeAreDevelopers World Congress North America wrapped up on Friday after two days absolutely crammed with content across eight (!) stages. There’s no way one person could absorb all of that in real time, but wonderfully, every session is recorded and every attendee got access to those recordings. So I went hunting for the best content across all eight stages across both days over the weekend, and summarized it all to the best of my ability, so you don’t have to wade through 237 talks.
It’s hard to come up with a narrative thread across such an enormous and topically diverse conference, but here’s my best attempt:
- Agents now write more code than humans can realistically review, making review the bottleneck, and people have a lot of strategies for doing that, not all of which work
- Software is simultaneously getting less deterministic, making production traces ever more important
- That non-deterministic software keeps escaping its test environments, making sandboxes mandatory
- Once you’ve figured out how to build agents safely, you have to make them well, and that has three parts:
- First you have to give them a context that works
- Second you have to make them able to effectively recall that context
- Thirdly you have to make them do all of that cost-effectively
- Once you’ve got a working agent, it will become somebody else’s user, and their agents will become yours. Agent Experience has become a top concern.
All of which is to say: software development has become mostly about AI. There were 237 sessions and I struggle to find one that wasn’t about AI in some way: building it, perfecting it and containing it. And that created the final and biggest theme, which is that: the job of the software developer has changed irrevocably, so how we train junior engineers has to change too.
I’ve done my best to gather up the best talks on each of those topics above and boil them down into the best practices common across speakers. That makes this post pretty long, but remember: 237 talks! To get the gist of the best, read on.
Build better agents with Arize
Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.
Prefer open source? Try Arize Phoenix for self-hosted, open source agent observability.
Code review is the bottleneck for AI-generated code
My own talk got upgraded last-minute to a main-stage keynote, which was very exciting (for me, I dunno about the audience). It was called The Death of the Code Review, and my main point was that human review has a hard speed limit, verified by real studies, and agent output doesn’t. The surprising thing is that agent-based code review has the same speed limit, also verified by data. So you have to get creative. I was by no means the only person to point this out. Aditya Jayaprakash of Blacksmith said the same thing in “The New Bottleneck in Software Development,” Prince Kohli said it on “The New Rules of Software Delivery,” and all 3 panelists on Gene Kim’s “Inside the AI-Native Engineering Org” (Amy Yuan of Snowflake, Siwei Shen of Coinbase, Hari Lingamagunta of Atlassian) said their orgs stall at review, not at coding.
Proposed solutions to this bottleneck were all over the conference. One way is to just not review it at all. My talk quoted Linear’s zero-review merge data, and Joseph Katsioloudes of Entire built an entire (pun intended) talk on it in “Review Was Already the Bottleneck. Then Agents Broke It Completely”: 38% of PRs now merge with no human review at all. But I (and others) concluded that doesn’t work: Manish Kapur of Sonar in “The reviewer can’t be the author” and Jeff An of Momentic on the “How to Trust Code You Didn’t Write” panel both said that a model reviewing its own output shares its own blind spots; An’s benchmarks found a leading coding agent caught only 15% of the bugs it introduced.
My proposal was to review at a higher level of abstraction, and I had smart people back me up on that. Kevin Lin of OpenAI in “The Agentic Engineering Loop” said he now judges only the scope, complexity and architecture fit of a PR. Colleen Lake of GitLab in “Agentic Drift” wants humans reviewing policies and exceptions rather than every change. Sam Jarvinen’s “The Autonomous Pull Request” reserved humans for intent, UX and unusual risk.
But if you move to a higher level of abstraction you have to introduce more safeguards on your production code.
Production evals and observability matter more as code review changes
If more stuff is getting into production without a human reading it, that means checks for “is it correct” also have to migrate from CI/CD tests to production, and that means evals, which is of course what we do at Arize AX. That was certainly the consensus at World Congress North America. Julia Furst Morgado of Dash0 in “Your Agents Need Observability Before They Need Better Models” called agents the biggest complexity spike since microservices, because they decide at runtime what code used to fix in advance. Him Raj Singh of PayPal in “Who Tests the AI?” laid out the core problem: the same input now produces different outputs, so classic testing is necessary and nowhere near sufficient.
The best practices that came out of this were remarkably consistent:
Instrument from day one with a shared vocabulary, and it’s OpenTelemetry. Furst Morgado pushed OpenTelemetry’s GenAI semantic conventions. The “Taming Rogue Agents” team (Anagha Rumade, Anjana Umapathy and Apoorva Jaiswal) used Arize’s own Phoenix to do it, also built on OpenTelemetry. Ashish Shubham’s “MicroAgents,” and the “Building Enterprise Software for AI Agents” panel all defaulted to OpenTelemetry traces.
Grade the path, not just the answer. “Taming Rogue Agents” opened with an agent that’s correct 95% of the time and still skips required policy checks, then scored goal, plan and action separately. Sachin Gupta in “Test Before Release, Enforce at Runtime” ran separate validators for policy, grounding and trajectory and found each one caught only its own failure type. Vivek Pandit of Turing in “Closing the Visibility Gap” recorded which assumptions were inferred and which tool call changed the plan, because in chip verification a missed bug costs millions.
Measure on live traffic, not just a golden set. Sofia Rest of Sentry in “21 Experiments in Six Weeks” used fix rate as a precision proxy and ran an experiment first to learn how much variance was noise. Milin Desai, also of Sentry, on “Running AI-Written Software in Production” told every team to build online and offline evals on top of their tracing data. Shruti Tiwari of Dell in “Where Judgment Went, and How to Put It Back” said to read raw transcripts, because they show what dashboards flatten.
Close the loop. Singh, Gupta, the “Taming Rogue Agents” team and Darko Mesaros of AWS in “I Don’t Trust AI Agents (And Neither Should You)” all landed on the same cycle: turn every production incident into a regression test, and rerun the suite on any model, prompt or tool change. That’s data you already have and are probably throwing away.
Contain the agent: the sandbox is the boundary
The timing here is good. NVIDIA announced the Open Agent Safety Platform this week, with OpenShell as the sandboxed runtime, and Justin Boitano previewed OpenShell on stage in his Acquired interview as a kernel-level boundary so probabilistic agents run inside deterministic limits. But NVIDIA were by no means the only people talking about the importance of sandboxes at World Congress.
Here’s the best practices I saw in common:
A real sandbox, not a container. Mark Cavage of Docker opened his keynote “Manufacturing trust: speed and safety in the age of agents” by having Claude escape a “contained” setup through the host Docker socket; containers isolate workloads, he said, but agents are actors. Rishab Kumar of Twilio in “rm -rf: Horror Stories From Unsandboxed AI Agents” and Ivan Burazin of Daytona in “Kubernetes Is Not Your Sandbox” both argued for microVMs built for the purpose.
Enforce at the tool call, because prompts are advice. Abdel Fane of OpenA2A in “Securing AI Agent Infrastructure,” Yoav Gal of Wonderful in “When Agents Became Users,” and the “New Security Stack” panel (Kara Sprague, Sagi Rodin, James Everingham) all said the same sentence in different words: rules in a system prompt are merely suggestions, but harness-level code restrictions are what really works.
Give the agent its own identity. Yoav Gal made agents first-class users at monday.com after a colleague was blamed for status changes her agent made overnight, because the audit log carried her name. Borko Djurkovic of Cohere in “Zero-Trust Architecture for Agentic AI” flowed user identity through every step with relationship-based access control. Tushar Jain of Docker in “Govern the Runtime, Not the Agent” agreed that agents shouldn’t impersonate the people who launched them.
Budgets and blast radius, not yes/no scopes. Sachin Malhotra of Anthropic in “Give the Agent a Budget, Not a[n API] Token” replayed an incident where an agent wiped about 200 workloads in under 2 minutes, and proposed rate limits on disruptive actions plus an “undo test.” Mesaros blocked refunds over $500 in the newly open-sourced Dogwood policy language. Suzanne Daniels of Microsoft in “Codifying Trade-offs” had the platform enforce constraints so the model is never trusted to. Notice that this is the same move as the code review one: you stop approving each action and start defining what’s reversible, how much is allowed, and who’s watching. Humans set the limits, agents work inside them, and the trace tells you whether it worked.
Build the agent: you (probably) don’t need frontier models
Perhaps not the most elegant segue between conference topics, but once you’ve figured out how to make your agent safe, you have to figure out how to make it good. There were two big piles of talks about how to do that: this first one is about selecting the right model, and a lot of that was about how you probably don’t need the most expensive model to make a good agent.
Some people argued the models barely matter. Bob Wambach titled his talk “It’s Not About the Models” and ranked models below architecture and data. Rahul Pandita of GitHub in “Lean Intelligence: Lessons from GitHub Copilot Data Science Efforts” showed a frontier model and a mid-tier model overlapping on 318 of 500 SWE-bench tasks, which is why a tiny classifier now routes each Copilot request.
The reason using a frontier model might not be the right choice is something I covered in Cost per successful task: the cost per token is not a good indicator of how much the model is going to cost you to complete the task, so you need to run evals on your specific workload to be sure. Manu Gurudatha of PagerDuty in “AI ROI: The Hard Unit Economics of AI-Native Engineering,” the “Unit Economics of AI” panel and Kelsey Hightower’s “ZTA: Zero Token Architecture” all argued for cost per accepted outcome as the number that matters.
The best practices on model selection were simple:
Route, don’t default. Pandita’s Hydra Fusion router came close to Opus at much lower cost per task. Viktoria Semaan of Databricks in “From Model Selection to Smart Routing” put all traffic behind a gateway with evals and budgets. Nitin Eusebius of AWS in “Edge AI” ran a 1B local orchestrator that escalates only the hard requests.
Plan expensive, execute cheap, verify strong. Wambach, Luis Pujols of GitHub in “Practices, Not Prompts,” and the “Building AI Products vs. Building With AI” panel (featuring Arize’s own Aparna Dhinakaran and Tamar Bercovici of Box) all gave the same recipe. And 5 unrelated talks reached for small zero-shot classifier models as the cheap tier: Li Yin, Ussama Baggili, Guillaume Lebedel, the PwC panel and Aparna.
Build evals so you can switch. The “Democratizing AI: Why Open Models Are Essential” panel (Joey Conway of NVIDIA, Olivier Lacombe of Google DeepMind, Varun Vontimitta of Meta) said to invest in evals so you can change providers without fear. Emmanuel Acheampong of Crusoe in “No Single Model to Rule Them All” said to hold no brand loyalty to models, harnesses or protocols. Boitano and Semaan independently put open models about 4 months behind the frontier, which makes switching cheap.
Context engineering: MCP tools and agent memory
In contrast to people arguing models don’t matter, everyone agreed that context engineering was crucial. MCP was the conference’s clearest success story. Ben Haefele of Webflow said most content development now runs through its MCP server. Jessica Garson Beauchemin of Runpod drove GPUs from Claude Code through about 75 MCP tools in “Managing GPUs by Just Asking.” Brendan Ittelson of Zoom said he’s fine with users never reopening Zoom’s UI because its interfaces are exposed over MCP.
Even the critics conceded the point. Vojta Kopal of Apify in “The MCP haters are half right” agreed MCP is the right remote interface and only complained about token cost. Guillaume Lebedel of StackOne in “Making (and Breaking) Agents by Adding 1,000 MCP Tools” watched tool schemas alone eat about 440k tokens, then fixed it with vector search over tools rather than abandoning the protocol.
How AI agents store and retrieve memory
Once you have context you need to recall it effectively, which made agent memory architectures another major topic. Emre Okcular of OpenAI in “Context Engineering: How machines remember and forget” and Carl Lapierre of Osedea in “Context Engineering Kung Fu” gave the same playbook: compact at a threshold, isolate sub-agents with fresh context, write memory outside the window.
Where that memory lives got 3 competing answers, all unsurprisingly recommending the thing their company is famous for. Elizabeth Fuentes Leone of AWS in “Vector, Graph, or Key Value?” said a plain SQL table is often enough. Guy Korland of FalkorDB in “Store Your AI Agent’s Memory in a Knowledge Graph” argued for temporal graphs that record when a fact stopped being true. Andrew Wong of Box in “File Systems Are the New Primitive for AI Agents” said models already know how to navigate a directory, so use one.
Agent memory security: Poisoning, conflicting instructions, and noise
Memory hygiene rules were consistent across all of them, though. Okcular named poisoning, conflicting instructions and noise as the 3 failure modes. Fuentes Leone warned that one poisoned graph node contaminates everything downstream of it. Ashok Prakash of Apple in “Responsible AI Architecture” scoped memory per session so poison can’t spread between users. Saloni Garg in “Red Teaming Your LLM App” showed a swapped email address persisting in memory as a “user preference.” Treat memory as an attack surface, because it is one.
Agent experience: Everyone else’s agent is your user
Once your agent works, assume everyone else’s does too, and they’ve already become your users. Han Wang of Mintlify on “Rethinking Developer Tools for the Agent Era” put agent traffic to Mintlify-hosted docs at about 67% and expects human traffic to become a rounding error.
It happened first in developer tools: there’s Haefele’s Webflow numbers I already mentioned, Stefan Lederer of Bitmovin on “Shipping AI Features People Actually Use” called agents first-class users of his products, and Monica Sarbu of Xata in “Databases in the Agent Era” said agents now create most new databases on some Postgres platforms.
Payments are also seeing agentic usage. Benjamin Smith of Stripe in “Your next customer won’t be human” showed shared payment tokens with spending caps, and Brian Whippo of the Algorand Foundation in “x402” noted that roughly 75 million HTTP 402 transactions in a recent month must be agents, because browsers don’t support the status code.
And so is the browser. Ajit Varma of Mozilla and Will Bryk of Exa on “The Browser Is Becoming an AI Runtime” expect AI searches to pass human searches by early 2027. Ankita Sood’s “WebMCP, A2UI, and Agent Experience” and Daniel Ostrovsky’s “Small LLM in Your Browser” both had web pages exposing their buttons as callable tools.
Unlike our other topics, no one is quite sure what best practices are for agents, yet. Some early suggestions include accurate docs served as markdown, an MCP server and an llms.txt file (Wang, Sam Bhagwat of Mastra, Matt Biilmann of Netlify in “From DX to AX”), a way for agents to sign up (Gil Feig of Merge saw nothing happen until Merge added hidden “if you’re an agent, start here” links), and tools that model the job rather than mirror your REST API (Haefele). Biilmann coined “agent experience” in early 2025, so the whole discipline is about 18 months old. Nobody says they have it figured out.
How AI is changing developer roles and junior hiring
All of this adds up to the argument I made in We are all Product Engineers now: agents are eating the software lifecycle from the bottom up, code generation first, review and operations next, and what’s left is figuring out what people actually want and defining what “good” means for a specific problem.
The “what’s left” part was the topic of more than one keynote. Lena Hall’s “Signal Layer” said the question is no longer whether something can be built but whether it should exist. The PwC panel “There’s No Such Thing as Vibe Strategy” said the bottleneck moved from capacity to judgment. Jenny Hwang of GitHub in “Going beyond the code” said the industry will be defined by deciding what’s worth building rather than how fast we type.
I said the job already exists under the name forward-deployed engineer, and Priyanka Vergadia in “The Reality of AI Adoption in Enterprises” and Tomislav Tipurić in “The Broken Rung” both described enterprises hiring exactly that role, with Tipurić adding that clients now push agencies toward outcome-based pricing.
I said the junior pipeline is broken, and the conference agreed. Tipurić cited steep declines in junior postings since 2023. The “Who Do We Hire Now?” panel (Becky Bucich, Ian Hughson, Jennifer Serrato, Kate Smeaton) admitted there’s no agreed replacement for the old junior-to-senior ladder. Brian Fox of Sonatype and Eli-Shaoul Khedouri of hCaptcha on “Defending at Machine Speed” asked how anyone builds judgment having only ever worked with AI.
And I said the fix is to rebuild the junior rung around product judgment rather than typing. The people actually hiring said the same thing. Matan Grinberg of Factory hires many juniors because agency now matters more than years of experience. Amy Yuan said Snowflake is still hiring new grads. Nikhil Gandhi of Yahoo Mail said juniors with domain guidance punch well above their level. Tipurić cited IBM reversing its junior hiring pause. Cassidy Williams of GitHub in “Our Brains in the AI Era” closed with 3 words: hire junior engineers.
Where do we go from here?
If you’ve made it through these 3000 words, congratulations, here’s what to take home:
- Move human review toward architecture, intent, and risk, and use production traces to investigate behavior and run evals.
- Put your agents in a real sandbox with their own identity and a budget.
- Pick models by cost per success and build the evals that let you switch.
- Spend your effort on context and memory instead.
- Build for the agents that are already reading your docs.
- And hire a junior this year, because the industry is going to need people who know what to build, and nobody’s training them yet.
See you all next year!