> ## Documentation Index
> Fetch the complete documentation index at: https://arize-ax.mintlify.site/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Find and Fix Issues in Your AI App with a Coding Agent

> Use a coding agent with Arize Skills and the AX CLI to trace your AI app, find failing traces, fix the code, and add an eval that stops the bug coming back.

<Card title="Sample app: Trailhead Outfitters support agent" icon="github" href="https://github.com/Arize-ai/tutorials/tree/main/python/cookbooks/find_and_fix_with_coding_agent">
  Clone this app to run the exact loop in this cookbook.
</Card>

If you already work through a coding agent, you can use it to improve your AI app too. Your agent can read your app's traces from Arize AX next to its source code, spot what's failing, fix it, and add an eval so the same problem can't come back. You review and approve each change.

The loop works with any coding agent that can load [Arize Skills](/docs/ax/skills/overview), such as Claude Code, Cursor, or Codex. Every step names the skill to use.

```mermaid theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
flowchart LR
  A[Instrument] --> B[Check]
  B --> C[Review]
  C --> D[Fix in code]
  C --> E[Guard with an eval]
  E --> F[Fix and prove]
  D --> G[Hand off]
  F --> G
  G --> C
```

## About the sample app

The sample app is a support agent for Trailhead Outfitters, a fictional outdoor gear store. It answers questions about orders, stock, and store policies using three tools, and runs on OpenAI or Anthropic. It ships without tracing so your agent can add it, and it has two bugs for your agent to find. The results below come from one run; your numbers will differ a little, because the model's answers vary.

## Before you start

You need:

* An [Arize AX account](https://app.arize.com/auth/join), with your API key and Space ID
* A coding agent with [Arize Skills installed](/docs/ax/skills/install), and the [AX CLI](/docs/api-clients/cli/overview) set up with an `ax` profile
* [uv](https://docs.astral.sh/uv/) and an OpenAI or Anthropic API key
* An [AI integration](/docs/ax/security-and-settings/integrations-playground/overview) in AX for the eval judge. If you don't have one, the `arize-ai-provider-integration` skill can create it.

## Set up the sample app

<Steps>
  <Step title="Clone the app">
    ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    git clone https://github.com/Arize-ai/tutorials.git
    cd tutorials/python/cookbooks/find_and_fix_with_coding_agent
    ```
  </Step>

  <Step title="Create the .env file">
    ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    cp .env.example .env
    ```

    Add your OpenAI or Anthropic key to the `.env` file. The app will use the relevant LLM based on the key you set.
  </Step>

  <Step title="Run the app">
    ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    uv run support.py "Do you have the Summit 2 Tent in stock?"
    ```
  </Step>
</Steps>

## Instrument the app and check the traces

<Steps>
  <Step title="Open the app in your coding agent" />

  <Step title="Instrument the app">
    Ask your agent:

    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-instrumentation skill to add Arize tracing to this app. Use the project name trailhead-support.
    ```

    The agent finds the OpenAI and Anthropic SDKs and a hand-written tool loop. It adds both provider instrumentors, and decorates the three tools and the `answer` function so each question becomes one trace: an agent span with LLM and tool spans under it.
  </Step>

  <Step title="Generate traces">
    Send 100 questions through the app:

    ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    uv run seed.py
    ```

    **In the sample app:** each trace has an `answer` agent span, `ChatCompletion` LLM spans, and a tool span for each tool call.
  </Step>

  <Step title="Check the traces are usable">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-instrumentation-health skill to check the instrumentation in the trailhead-support project.
    ```

    **In the sample app:** healthy. Every root is an agent span with an input and output, every LLM span has token counts, and no spans are orphaned.
  </Step>
</Steps>

## Review the traces

<Steps>
  <Step title="Ask your agent to review the traces">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-trace skill to look at the last 7 days of traces for this app and tell me what's failing.
    ```

    The agent exports the spans, groups the traces by what the tools returned and how the agent answered, and reads the app's code to work out why.
  </Step>

  <Step title="Read the findings">
    **In the sample app**, it found both bugs:

    * **Order lookups fail (22 of 39 order questions).** The model doesn't know today's date, and `lookup_order` needs an exact date. For "yesterday" or "last week" the agent asks the customer for the date. For "September 23" it guesses the wrong year (`2023-09-23`) and reports that the order doesn't exist.
    * **Policy answers are made up (about 33 of 40 policy questions).** `get_policy` is defined but never offered to the model, so it answers from general knowledge: 30-day returns instead of 21, sale items returnable when they're final sale, and no restocking fee when electronics have a 10% fee.

    Nothing in these traces has an error status, and the latency looks normal. The agent found the problems by comparing what the app said with what its tools returned and with the policies in `store.py`.
  </Step>
</Steps>

## Fix the order lookups in code

This one is a plain code bug, so fix it directly.

<Steps>
  <Step title="Fix the bug">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Fix the order lookup problem.
    ```

    **In the sample app:** the agent adds today's date to the system prompt and makes `order_date` optional, so `lookup_order` can list a customer's recent orders.
  </Step>

  <Step title="Generate new traces">
    Send another 100 questions through the fixed app:

    ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    uv run seed.py
    ```
  </Step>

  <Step title="Check the fix">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-trace skill to check the order lookups in the new traces sent since the fix.
    ```

    **In the sample app:** on the new traffic, **40 of 40 order questions resolve**.
  </Step>
</Steps>

## Guard the policy answers with an eval

You could fix the policy bug the same way. But wrong answers like these can come back with any prompt or model change, and the real policy usually lives outside the repo, so add an eval that catches them first.

<Steps>
  <Step title="Create the evaluator">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-evaluator skill to create an LLM-as-judge evaluator that checks policy answers against our store policy in store.py. Just create the evaluator for now, don't run it yet.
    ```

    The agent creates an evaluator with the policy text as the reference.
  </Step>

  <Step title="Run it on recent traffic">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-evaluator skill to run the evaluator on the last day of agent spans in trailhead-support.
    ```

    The agent creates a task that scores agent spans, with the customer's question as the input and the agent's reply as the output.

    **In the sample app:** the eval marks **37 of 48 policy answers as inaccurate**, which matches what the review found. The judge's explanations name each mistake, for example "the store policy allows returns within 21 days of delivery".
  </Step>
</Steps>

## Fix the policy answers and prove it

<Steps>
  <Step title="Create a dataset of the failures">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-dataset skill to create a dataset from the traces the Policy Accuracy eval marked inaccurate.
    ```

    **In the sample app:** the dataset has 9 distinct failing questions.
  </Step>

  <Step title="Run a baseline experiment">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-experiment skill to run the dataset through the current app as a baseline experiment, and score it with the Policy Accuracy evaluator.
    ```

    **In the sample app:** the baseline scores around 3%.
  </Step>

  <Step title="Fix the bug">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Fix the policy answers bug. Then use the arize-experiment skill to run the dataset through the fixed app as a second experiment, score it with the Policy Accuracy evaluator, and compare it with the baseline.
    ```

    **In the sample app:** the agent offers `get_policy` to the model and adds one line to the system prompt: answer policy questions only from what `get_policy` returns. The fixed app should now score 100% in the experiment.
  </Step>
</Steps>

## Keep the eval running

The task you used to score recent traffic ran once. Turn it into an online eval, so it scores new traffic as it arrives and the score drops if a later change brings the bug back.

<Steps>
  <Step title="Turn on online evaluation">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-evaluator skill to update the Policy Accuracy task in trailhead-support so it runs continuously on new agent spans.
    ```

    The agent updates the task to run continuously on new agent spans.
  </Step>

  <Step title="Generate new traces">
    Send another 100 questions through the fixed app:

    ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    uv run seed.py
    ```

    New spans can take a few minutes to be scored.
  </Step>

  <Step title="Check the new scores">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Use the arize-trace skill to check the Policy Accuracy scores on the agent spans sent since the continuous task was enabled.
    ```

    **In the sample app:** after both fixes, all the policy answers in the new traffic scored accurate.
  </Step>
</Steps>

## Hand off for review

<Steps>
  <Step title="Write the pull request description">
    ```text wrap theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
    Summarize both fixes with their before and after numbers, and use the arize-link skill to add links to the experiments and the evaluator. Write it as the description you would use for a pull request. Don't create a branch, commit, push, or open a pull request: this app is part of the Arize tutorials repo.
    ```

    The agent writes the description, with the evidence and links for each fix.
  </Step>

  <Step title="Review the summary">
    Check that each fix has its before and after numbers, and open the links to see the experiments and evaluator in Arize AX.
  </Step>
</Steps>

<Note>
  In your own app, ask your agent to open the pull request with this description, then review the pull request and decide whether to merge.
</Note>

## Keep the loop going

* Run the trace review on a schedule, for example weekly, with the same prompt.
* Turn your review prompt into a skill once it works well, so it runs the same way every time. Save the prompt and the grouping you want as a `SKILL.md`; see [Skills](/docs/ax/skills/overview) for how skills work.
* Turn on [Signal](/docs/ax/observe/signal) to have Arize AX surface ranked issues from production traces automatically.

## Next steps

<CardGroup cols={3}>
  <Card title="Skill catalog" icon="list" href="/docs/ax/skills/catalog">
    Every Arize skill and what it does.
  </Card>

  <Card title="AX CLI" icon="terminal" href="/docs/api-clients/cli/overview">
    The full `ax` command reference.
  </Card>

  <Card title="Create evaluators" icon="scale-balanced" href="/docs/ax/evaluate/create-evaluators">
    Design evaluators that catch the failures you care about.
  </Card>
</CardGroup>
