For years, software teams have had a well-understood path from development to production.
They write code. They test it. They move it through sandbox, staging, and production environments. They use CI/CD, monitoring, and release processes to catch regressions before they affect customers.
Rahul Todkar, VP of Data and AI at Tripadvisor, believes AI products need the same level of rigor.
As Tripadvisor builds more AI-powered travel experiences, the challenge is no longer just building models that work in isolation. The company is operating a growing mix of recommendation models, search models, pricing models, predictive systems, LLM-powered products, multimodal models, and agentic workflows.
That makes the production question much harder.
How do you monitor both traditional ML and LLM systems in one place? How do you trace and debug multi-agent workflows before they fail? How do you know whether a production AI experience is working as intended for travelers at scale?
For Tripadvisor, Arize has become part of the answer.
“From a measurements standpoint, there’s a fantastic platform and features that have been built in Arize, which our teams don’t have to necessarily create on our own,” Todkar says. “Comprehensive visibility and depth of measurement are the two big reasons where we are seeing benefits from Arize today.”
The challenge: AI products needed a production lifecycle as rigorous as software
Tripadvisor has a number of agents built into it’s products, and that range creates a need for maintaining standards amid potential sprawl.
Some power the company’s conversational chat interface, helping combine recommendations, search, personalization, and contextual insights into customer-facing experiences across the website and app. Others are externalized through MCPs and other mechanisms, helping customers outside the Tripadvisor site.
That shift changes what production readiness means.
A traditional software feature can move through an established development lifecycle. AI systems, especially LLM-based and agentic systems, are still catching up to that level of maturity.
“You have a certain software development lifecycle, an SDLC,” Todkar says. “You build software, you test it, you have a sandbox, a staging environment, and a production environment. You have to go through that process. As you look at AI products and AI applications, we have not fully developed and matured that process.”
For Todkar, moving from experimentation to production requires more than a promising demo or a working prototype. It requires the same kind of process orientation that engineering teams expect from production software.
“Moving from experimentation to production has to have that level of rigor and process orientation towards taking these systems into real customer-facing products.”
That was the core limitation Tripadvisor needed to solve: its AI stack was expanding quickly, but the operational practices around AI needed to mature just as quickly.
The shift: observability for both traditional ML and LLM-based systems
Tripadvisor’s AI product development lifecycle now spans two worlds.
On one side are the traditional models that have powered digital products for years: recommendations, search, auction models, pricing systems, and predictive models.
And on the other side are newer LLM-based systems: conversational interfaces, multimodal models, agentic workflows, and AI products that reason across multiple steps.
Many companies treat these as separate observability problems. One set of tools monitors traditional ML while another monitors LLM applications. And another traces agent workflows—while yet another handles evals.
Todkar wanted to avoid that fragmentation.
“What we want is a partner like Arize, which offers monitoring and observability for both of these AI worlds — LLM models as well as non-LLM models,” he says.
That breadth mattered because Tripadvisor’s production AI environment was not moving from old ML to new AI. It was combining both.
The company still needed visibility into the performance of established ML systems. But it also needed to understand how LLM-based products behaved in production, how agentic workflows executed, and how failures could be found before they reached travelers.
“We don’t have to have multiple point solutions just for model observability and model monitoring,” Todkar explains. “You can do that all together with Arize.”
Why traces and evals matter for production agents
As Tripadvisor began putting more LLM-based products into production, one issue became increasingly important: understanding agentic workflows.
With a traditional model, the unit of analysis might be a prediction, ranking, score, or classification. With an agent, the unit of analysis is often a workflow: a sequence of tool calls, model responses, retrieval steps, decisions, and handoffs.
That makes debugging harder.
A failure may not come from a single bad output. It may come from latency in one step, bad data feeding into the model, an unexpected response from a tool, or an agent path that looks reasonable locally but fails across the full workflow.
Todkar explains this: “As we started putting LLM-based products into production, one thing that started to be hard was: how do you understand the agentic workflows, trace them, and debug them in production before certain things fail?”
This is where Arize helped Tripadvisor move from surface-level monitoring to deeper production measurement.
The team could ingest traces, evaluate agent behavior, and inspect where production systems were not performing as expected. Instead of waiting for failures to become customer-visible, teams could catch issues earlier in the development and deployment process.
“As we looked at Arize, we could catch these bugs ahead of time before they went into production,” Todkar says.
The application: Tripadvisor’s AI trip planner
Todkar pointed to Tripadvisor’s AI trip planner as one of the company’s biggest AI applications.
The product uses multi-agent systems in production to help customers plan travel experiences at scale. That makes observability and evals especially important.
A trip planner is not a simple one-turn interaction. A traveler may have preferences, constraints, destinations, budgets, activities, timing, and context that all affect the quality of the recommendation. The system has to reason across multiple steps and deliver an experience that feels personalized and relevant.
In that kind of environment, production issues can come from many places.
Todkar described examples where the team was able to catch models not working as intended, latency issues, and cases where the data feeding into the system was not up to the expected quality bar.
“As we looked at building those features, there were many instances where we were able to catch models not working, latency issues, or data feeding into them that wasn’t up to par,” he says. “All those things we have been able to catch much better with Arize now.”
That is the operational value of deeper measurement: not just seeing that something failed, but understanding which part of the AI workflow needs attention.
For a production travel assistant, that difference matters. The customer experience depends on the entire system working together: models, data, orchestration, latency, and agent behavior.
The process change: building toward automated evaluation and governance
Todkar is clear that Tripadvisor is still early in building the full process around continuous agent evaluation. But the direction is clear.
As the company builds more agentic workflows, it wants these systems to become as automated as possible while still maintaining robust testing, monitoring, and governance.
“The moment to keep on evaluating agents is an interesting one,” Todkar says. “We are still at the early stage of building that process and the robustness around that.”
That means evals are not just a one-time pre-launch check. They become part of a broader production discipline: testing agentic workflows, monitoring behavior, identifying failures, and governing what agents are allowed to do.
“As we build and think more and more about agentic workflows and agentic commerce, we are trying to take that into a way where these systems are as automated as possible,” Todkar explains.
For Tripadvisor, this is part of the larger shift from AI experimentation to production AI operations.
The question is not simply whether a model performs well in a notebook or whether an agent can complete a demo. It’s whether the full system can be tested, monitored, governed, and improved once real customers are using it.
Why Arize: one production platform for AI observability, evals, and governance
When asked whether he would recommend Arize to other teams building AI products, Todkar was direct.
“Absolutely yes,” he says. “For anybody who is building AI products or AI solutions, Arize is a must for your production system.”
He called out two reasons:
- Arize provides a comprehensive model monitoring, observability, and governance platform that covers both traditional ML models and LLM-based systems.
- Arize is building for the newest production challenges around AI agents and agentic workflows.
Todkar explains it this way: “It is a comprehensive model monitoring, observability, and governance platform which takes into account traditional models as well as LLM-based models — a comprehensive suite in one place.”
That “one place” point is important.
For teams operating complex AI systems, fragmented tooling can create fragmented understanding. The more systems involved in production, the more important it becomes to connect signals across models, data, traces, evals, latency, and customer-facing behavior.
Arize gave Tripadvisor a way to bring more of that production AI lifecycle into a shared measurement layer.
“For anybody who is building AI products or AI solutions, Arize is a must for your production system.”
What comes next: agentic commerce
Todkar sees agentic systems as part of a broader platform shift.
Just as web, mobile, and cloud changed how users interacted with technology, he believes agents will change the interface between travelers, digital products, and commerce.
In travel, that future points toward agentic commerce: systems that can act on a traveler’s behalf.
A traveler may define their preferences ahead of time. An agent could then interact with multiple providers, reason across options, and help execute parts of the travel planning or booking process.
But that future also raises the operational bar.
If agents are going to act on behalf of customers, companies need more than model performance. They need deployment support, governance, monitoring, observability, policy enforcement, infrastructure, and orchestration.
That is why Todkar sees production AI infrastructure as strategic.
For Tripadvisor, the path to agentic commerce is not just about building more capable agents. It is about building the systems that make those agents reliable, measurable, and governable in production.
“How do you take these systems into production?” he asks. “What is the right level of production deployment support? How do you get the right level of governance, monitoring, observability, and policy enforcement? That becomes really important.”
The challenge: AI products needed a production lifecycle as rigorous as software
Tripadvisor has different types of agents running in production today.
The shift: observability for both traditional ML and LLM-based systems
Tripadvisor’s AI product development lifecycle now spans two worlds.
Why traces and evals matter for production agents
As Tripadvisor began putting more LLM-based products into production, one issue became increasingly important: understanding agentic workflows .
The application: Tripadvisor’s AI trip planner
Todkar pointed to Tripadvisor’s AI trip planner as one of the company’s biggest AI applications.
The process change: building toward automated evaluation and governance
Todkar is clear that Tripadvisor is still early in building the full process around continuous agent evaluation. But the direction is clear.
Why Arize: one production platform for AI observability, evals, and governance
When asked whether he would recommend Arize to other teams building AI products, Todkar was direct.
What comes next: agentic commerce
Todkar sees agentic systems as part of a broader platform shift.