Sample app: Trailhead Outfitters support agent
About the sample app
The sample app is a support agent for Trailhead Outfitters, a fictional outdoor gear store. It answers questions about orders, stock, and store policies using three tools, and runs on OpenAI or Anthropic. It ships without tracing so your agent can add it, and it has two bugs for your agent to find. The results below come from one run; your numbers will differ a little, because the model’s answers vary.Before you start
You need:- An Arize AX account, with your API key and Space ID
- A coding agent with Arize Skills installed, and the AX CLI set up with an
axprofile - uv and an OpenAI or Anthropic API key
- An AI integration in AX for the eval judge. If you don’t have one, the
arize-ai-provider-integrationskill can create it.
Set up the sample app
Clone the app
Create the .env file
.env file. The app will use the relevant LLM based on the key you set.Run the app
Instrument the app and check the traces
Open the app in your coding agent
Instrument the app
answer function so each question becomes one trace: an agent span with LLM and tool spans under it.Generate traces
answer agent span, ChatCompletion LLM spans, and a tool span for each tool call.Check the traces are usable
Review the traces
Ask your agent to review the traces
Read the findings
- Order lookups fail (22 of 39 order questions). The model doesn’t know today’s date, and
lookup_orderneeds an exact date. For “yesterday” or “last week” the agent asks the customer for the date. For “September 23” it guesses the wrong year (2023-09-23) and reports that the order doesn’t exist. - Policy answers are made up (about 33 of 40 policy questions).
get_policyis defined but never offered to the model, so it answers from general knowledge: 30-day returns instead of 21, sale items returnable when they’re final sale, and no restocking fee when electronics have a 10% fee.
store.py.Fix the order lookups in code
This one is a plain code bug, so fix it directly.Fix the bug
order_date optional, so lookup_order can list a customer’s recent orders.Generate new traces
Check the fix
Guard the policy answers with an eval
You could fix the policy bug the same way. But wrong answers like these can come back with any prompt or model change, and the real policy usually lives outside the repo, so add an eval that catches them first.Create the evaluator
Run it on recent traffic
Fix the policy answers and prove it
Create a dataset of the failures
Run a baseline experiment
Fix the bug
get_policy to the model and adds one line to the system prompt: answer policy questions only from what get_policy returns. The fixed app should now score 100% in the experiment.Keep the eval running
The task you used to score recent traffic ran once. Turn it into an online eval, so it scores new traffic as it arrives and the score drops if a later change brings the bug back.Turn on online evaluation
Generate new traces
Check the new scores
Hand off for review
Write the pull request description
Review the summary
Keep the loop going
- Run the trace review on a schedule, for example weekly, with the same prompt.
- Turn your review prompt into a skill once it works well, so it runs the same way every time. Save the prompt and the grouping you want as a
SKILL.md; see Skills for how skills work. - Turn on Signal to have Arize AX surface ranked issues from production traces automatically.
Next steps
Skill catalog
AX CLI
ax command reference.