How LG Uplus is building better AI customer service agents with evaluation-driven development

How LG Uplus is building better AI customer service agents with evaluation-driven development

Published August 17 2026 6 min read

AI has become an important part of how LG Uplus serves its customers.

The company is building AI across its contact center, helping handle customer conversations, support agents during calls, and automate work after each interaction. With more than 30 million subscribers, even small improvements can make a meaningful difference.

As those systems became more sophisticated, the team found that building the agent was only part of the challenge. The harder question was how to continuously evaluate and improve it once it was in use.

That realization changed the way the team develops AI.

Today, evaluation is built into the development process itself. Using Arize AX, LG Uplus created a workflow that combines production traces, user feedback, and domain expertise to continuously improve its AI agents.

Building AI across the customer journey

The AI Contact Center team isn’t focused on a single chatbot or assistant. Their goal is to support the full customer experience. That includes helping route and process customer calls, assisting human agents during conversations, and reducing the manual work that happens after a call ends.

“We’re building AI that handles the beginning and end of the customer journey — from processing calls in the contact center, helping customer service representatives, and supporting post-call work” describes Minkyu Ha, a senior software engineering manager at LG Uplus.

Because those systems touch real customer interactions, quality matters just as much as capability.

Customer service doesn’t have one right answer

One of the biggest lessons for the team has been that evaluating customer service AI is fundamentally different from evaluating many other AI applications.

In customer conversations, there often isn’t a single correct response. Success depends on context, judgment, and the experience the customer has. That makes evaluation much harder than simply checking whether an answer matches a reference response.

As Minkyu Ha explains: “Not everything can be quantified. When it can’t be quantified, the correct answer becomes ambiguous.”

The team experimented with different approaches, including random sampling and reviewing specific cases. Over time, they found that the most valuable signal came from the people actually using the system.

“There was nothing better than user feedback,” Minkyu Ha says. “The highest-quality feedback came from end users and from domain experts.”

Customer comments helped identify where the AI struggled. Contact center specialists provided the context needed to judge conversations that couldn’t be measured with a simple metric.

Building evaluation into the development process

For LG Uplus, one of the biggest changes was changing how the team builds AI.

Rather than developing an agent first and thinking about evaluation later, the team now starts by building evaluation datasets and continuously refining them as the system improves.

Minkyu Ha says that shift changed the team’s entire development cycle.

“The development process itself changed,” he explains. “We adopted an evaluation-driven development approach with Arize and continuously improved performance by building evaluation datasets.”

Production traces flow into the team’s internal systems through APIs, creating what Ha describes as a closed-loop pipeline for continuous improvement.

Instead of manually collecting traces, writing additional code, and running separate evaluation cycles, much of that process now happens automatically.

“Once the traces are collected, the evaluation cycle becomes automated,” Minkyu Ha says. “That development time simply disappears.”

The result is a workflow where learning from production happens much faster, giving engineers more time to focus on improving the agent itself.

Looking beyond the final response

As AI systems become more agentic, understanding the final answer is no longer enough.The intermediate steps matter too. Which tools did the agent choose? How did it route the request? Where did the workflow break down?

Those questions are often difficult to answer by looking only at outputs.

According to Minkyu Ha, that visibility has become one of the biggest improvements since introducing Arize. He explains: “The most important part of an agent flow is selecting tools and optimizing routing paths. Before, evaluating those intermediate processes was difficult. After introducing Arize, that improved significantly.”

Being able to inspect those intermediate decisions helps the team understand why an interaction succeeded or failed, making it easier to improve future versions of the agent.

AI hasn’t replaced domain experts. It’s made them more important.

One of the most interesting observations from Minkyu Ha is that better AI hasn’t reduced the need for human expertise. It’s the opposite.

Because many customer service interactions don’t have an objectively correct answer, experienced reviewers still play a critical role in evaluation.

At LG Uplus, contact center knowledge experts review conversations that can’t be judged automatically.

Minkyu Ha believes their role as an AI engineering team has only become more important as AI systems become more capable. “Paradoxically, now that the age of AI agents has arrived, the role of domain experts has become even bigger,” he says.

It’s a reminder that building reliable AI isn’t just a technical challenge. It also requires the people who understand the business, the customers, and what a successful interaction actually looks like.

Building a culture around evaluation

After attending Arize:Observe, Minkyu Ha saw that many of the industry’s leading AI teams were moving in a similar direction.

Evaluation was no longer treated as a final quality check. It had become part of the development process itself.

For LG Uplus, that’s the direction the company will continue investing in.

Minkyu Ha believes building an evaluation-driven culture requires more than engineering teams. Domain experts, business stakeholders, and product teams all need to participate in defining and improving AI quality.

“I think this is an essential process not only for development teams, but also for domain experts and business stakeholders across the company,” he says.

As AI systems continue to evolve, LG Uplus sees evaluation as more than a way to measure quality. It’s becoming the foundation for building AI that can continuously learn, improve, and better serve millions of customers.

Building AI across the customer journey

The AI Contact Center team isn’t focused on a single chatbot or assistant. Their goal is to support the full customer experience. That includes helping route and process customer calls, assisting human agents during conversations, and reducing the manual work that happens after a call ends.

Customer service doesn't have one right answer

One of the biggest lessons for the team has been that evaluating customer service AI is fundamentally different from evaluating many other AI applications.

Building evaluation into the development process

For LG Uplus, one of the biggest changes was changing how the team builds AI.

Looking beyond the final response

As AI systems become more agentic , understanding the final answer is no longer enough.The intermediate steps matter too. Which tools did the agent choose? How did it route the request? Where did the workflow break down?

AI hasn't replaced domain experts. It's made them more important.

One of the most interesting observations from Minkyu Ha is that better AI hasn’t reduced the need for human expertise. It’s the opposite.

Building a culture around evaluation

After attending Arize:Observe, Minkyu Ha saw that many of the industry’s leading AI teams were moving in a similar direction.

Don’t ship vibes.

Arize gives AI teams observability and evals to understand and improve agent performance.