Skip to main content

Sample app: Trailhead Outfitters support agent

Clone this app to run the exact loop in this cookbook.
If you already work through a coding agent, you can use it to improve your AI app too. Your agent can read your app’s traces from Arize AX next to its source code, spot what’s failing, fix it, and add an eval so the same problem can’t come back. You review and approve each change. The loop works with any coding agent that can load Arize Skills, such as Claude Code, Cursor, or Codex. Every step names the skill to use.

About the sample app

The sample app is a support agent for Trailhead Outfitters, a fictional outdoor gear store. It answers questions about orders, stock, and store policies using three tools, and runs on OpenAI or Anthropic. It ships without tracing so your agent can add it, and it has two bugs for your agent to find. The results below come from one run; your numbers will differ a little, because the model’s answers vary.

Before you start

You need:
  • An Arize AX account, with your API key and Space ID
  • A coding agent with Arize Skills installed, and the AX CLI set up with an ax profile
  • uv and an OpenAI or Anthropic API key
  • An AI integration in AX for the eval judge. If you don’t have one, the arize-ai-provider-integration skill can create it.

Set up the sample app

1

Clone the app

2

Create the .env file

Add your OpenAI or Anthropic key to the .env file. The app will use the relevant LLM based on the key you set.
3

Run the app

Instrument the app and check the traces

1

Open the app in your coding agent

2

Instrument the app

Ask your agent:
The agent finds the OpenAI and Anthropic SDKs and a hand-written tool loop. It adds both provider instrumentors, and decorates the three tools and the answer function so each question becomes one trace: an agent span with LLM and tool spans under it.
3

Generate traces

Send 100 questions through the app:
In the sample app: each trace has an answer agent span, ChatCompletion LLM spans, and a tool span for each tool call.
4

Check the traces are usable

In the sample app: healthy. Every root is an agent span with an input and output, every LLM span has token counts, and no spans are orphaned.

Review the traces

1

Ask your agent to review the traces

The agent exports the spans, groups the traces by what the tools returned and how the agent answered, and reads the app’s code to work out why.
2

Read the findings

In the sample app, it found both bugs:
  • Order lookups fail (22 of 39 order questions). The model doesn’t know today’s date, and lookup_order needs an exact date. For “yesterday” or “last week” the agent asks the customer for the date. For “September 23” it guesses the wrong year (2023-09-23) and reports that the order doesn’t exist.
  • Policy answers are made up (about 33 of 40 policy questions). get_policy is defined but never offered to the model, so it answers from general knowledge: 30-day returns instead of 21, sale items returnable when they’re final sale, and no restocking fee when electronics have a 10% fee.
Nothing in these traces has an error status, and the latency looks normal. The agent found the problems by comparing what the app said with what its tools returned and with the policies in store.py.

Fix the order lookups in code

This one is a plain code bug, so fix it directly.
1

Fix the bug

In the sample app: the agent adds today’s date to the system prompt and makes order_date optional, so lookup_order can list a customer’s recent orders.
2

Generate new traces

Send another 100 questions through the fixed app:
3

Check the fix

In the sample app: on the new traffic, 40 of 40 order questions resolve.

Guard the policy answers with an eval

You could fix the policy bug the same way. But wrong answers like these can come back with any prompt or model change, and the real policy usually lives outside the repo, so add an eval that catches them first.
1

Create the evaluator

The agent creates an evaluator with the policy text as the reference.
2

Run it on recent traffic

The agent creates a task that scores agent spans, with the customer’s question as the input and the agent’s reply as the output.In the sample app: the eval marks 37 of 48 policy answers as inaccurate, which matches what the review found. The judge’s explanations name each mistake, for example “the store policy allows returns within 21 days of delivery”.

Fix the policy answers and prove it

1

Create a dataset of the failures

In the sample app: the dataset has 9 distinct failing questions.
2

Run a baseline experiment

In the sample app: the baseline scores around 3%.
3

Fix the bug

In the sample app: the agent offers get_policy to the model and adds one line to the system prompt: answer policy questions only from what get_policy returns. The fixed app should now score 100% in the experiment.

Keep the eval running

The task you used to score recent traffic ran once. Turn it into an online eval, so it scores new traffic as it arrives and the score drops if a later change brings the bug back.
1

Turn on online evaluation

The agent updates the task to run continuously on new agent spans.
2

Generate new traces

Send another 100 questions through the fixed app:
New spans can take a few minutes to be scored.
3

Check the new scores

In the sample app: after both fixes, all the policy answers in the new traffic scored accurate.

Hand off for review

1

Write the pull request description

The agent writes the description, with the evidence and links for each fix.
2

Review the summary

Check that each fix has its before and after numbers, and open the links to see the experiments and evaluator in Arize AX.
In your own app, ask your agent to open the pull request with this description, then review the pull request and decide whether to merge.

Keep the loop going

  • Run the trace review on a schedule, for example weekly, with the same prompt.
  • Turn your review prompt into a skill once it works well, so it runs the same way every time. Save the prompt and the grouping you want as a SKILL.md; see Skills for how skills work.
  • Turn on Signal to have Arize AX surface ranked issues from production traces automatically.

Next steps

Skill catalog

Every Arize skill and what it does.

AX CLI

The full ax command reference.

Create evaluators

Design evaluators that catch the failures you care about.