> ## Documentation Index
> Fetch the complete documentation index at: https://arizeai-433a7140.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluate a Prompt Change in Code

> Build the dataset, task, evaluator, and experiment with the SDK, then compare a baseline prompt against an edited one.

[Evaluate](/docs/phoenix/get-started/get-started-evaluations) scores two prompts with an LLM judge in one short script. This page is the longer code walkthrough: a deterministic code evaluator instead of a judge, examples with ids and metadata, and a dataset you built from traces in the UI. The examples are billing-support replies. It needs a running Phoenix and no model key.

You define two tasks, a baseline and an edited prompt, score both with one deterministic evaluator, and read the difference. The code lives in [examples/quickstarts](https://github.com/Arize-ai/phoenix/tree/main/examples/quickstarts), ready to run.

## Build and run the experiment

The script creates its own dataset, so you can run it without traces in Phoenix. Choose Python or TypeScript in each step.

<Steps>
  <Step title={<span className="step-title">Install experiment packages</span>}>
    <Tabs>
      <Tab title="Python">
        ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        pip install -U arize-phoenix-client
        ```
      </Tab>

      <Tab title="TypeScript">
        ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        npm install @arizeai/phoenix-client tsx
        ```
      </Tab>
    </Tabs>
  </Step>

  <Step title={<span className="step-title">Create an experiment file</span>}>
    The dataset uses two billing support examples; replace them with selected traces or examples from your agent when you are ready.

    <Tabs>
      <Tab title="Python">
        Create a file named `experiment.py`.

        ```python theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        from phoenix.client import Client

        client = Client()

        dataset = client.datasets.create_dataset(
            name="support-quickstart",
            dataset_description="Support reply examples from the Phoenix quickstart.",
            examples=[
                {
                    "id": "invoice-change",
                    "input": {"text": "Draft a short reply about why an invoice changed."},
                    "output": {
                        "text": "Explain that invoice totals can change when usage, taxes, or plan adjustments are applied."
                    },
                    "metadata": {"kind": "billing"},
                },
                {
                    "id": "renewal-plan-change",
                    "input": {
                        "text": "A customer says their renewal invoice is higher after changing plans."
                    },
                    "output": {
                        "text": "Explain that plan changes can cause prorated usage charges and updated taxes."
                    },
                    "metadata": {"kind": "billing"},
                },
            ],
        )


        def baseline_task(input):
            if "renewal" in input["text"]:
                return "Renewal amounts can change after account updates."
            return "Invoice totals can change when billing settings change."


        def edited_task(input):
            if "renewal" in input["text"]:
                return (
                    "Plan changes can create prorated usage charges and updated taxes on a renewal invoice."
                )
            return "Invoice totals can change when usage, taxes, or plan adjustments are applied."


        def covers_billing_context(output):
            text = output.lower()
            has_usage = "usage" in text or "prorat" in text
            has_taxes = "tax" in text
            has_plan = "plan" in text
            return has_usage and has_taxes and has_plan


        client.experiments.run_experiment(
            dataset=dataset,
            task=baseline_task,
            evaluators=[covers_billing_context],
            experiment_name="support-response-baseline",
        )

        client.experiments.run_experiment(
            dataset=dataset,
            task=edited_task,
            evaluators=[covers_billing_context],
            experiment_name="support-response-edited",
        )
        ```
      </Tab>

      <Tab title="TypeScript">
        Create a file named `experiment.ts`.

        ```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        import { createClient } from "@arizeai/phoenix-client";
        import { createDataset } from "@arizeai/phoenix-client/datasets";
        import {
          asExperimentEvaluator,
          runExperiment,
        } from "@arizeai/phoenix-client/experiments";
        import type { Example } from "@arizeai/phoenix-client/types/datasets";

        const client = createClient();

        const dataset = await createDataset({
          client,
          name: "support-quickstart",
          description: "Support reply examples from the Phoenix quickstart.",
          examples: [
            {
              id: "invoice-change",
              input: { text: "Draft a short reply about why an invoice changed." },
              output: {
                text: "Explain that invoice totals can change when usage, taxes, or plan adjustments are applied.",
              },
              metadata: { kind: "billing" },
            },
            {
              id: "renewal-plan-change",
              input: {
                text: "A customer says their renewal invoice is higher after changing plans.",
              },
              output: {
                text: "Explain that plan changes can cause prorated usage charges and updated taxes.",
              },
              metadata: { kind: "billing" },
            },
          ],
        });

        const baselineTask = (example: Example): string => {
          const input = example.input as { text: string };

          if (input.text.includes("renewal")) {
            return "Renewal amounts can change after account updates.";
          }
          return "Invoice totals can change when billing settings change.";
        };

        const editedTask = (example: Example): string => {
          const input = example.input as { text: string };

          if (input.text.includes("renewal")) {
            return "Plan changes can create prorated usage charges and updated taxes on a renewal invoice.";
          }
          return "Invoice totals can change when usage, taxes, or plan adjustments are applied.";
        };

        const coversBillingContext = asExperimentEvaluator({
          name: "covers_billing_context",
          kind: "CODE",
          evaluate: ({ output }) => {
            const text = String(output).toLowerCase();
            const hasUsage = text.includes("usage") || text.includes("prorat");
            const hasTaxes = text.includes("tax");
            const hasPlan = text.includes("plan");
            const passes = hasUsage && hasTaxes && hasPlan;
            return {
              label: passes ? "True" : "False",
              score: passes ? 1 : 0,
            };
          },
        });

        await runExperiment({
          client,
          dataset,
          task: baselineTask,
          evaluators: [coversBillingContext],
          experimentName: "support-response-baseline",
        });

        await runExperiment({
          client,
          dataset,
          task: editedTask,
          evaluators: [coversBillingContext],
          experimentName: "support-response-edited",
        });
        ```
      </Tab>
    </Tabs>

    Each task takes one dataset example and returns an agent output. Swap in your agent's entry point. The evaluator is intentionally deterministic: it checks for the billing context the response must include before you add broader LLM-as-a-judge evals.
  </Step>

  <Step title={<span className="step-title">Run the experiment</span>}>
    <Tabs>
      <Tab title="Python">
        ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        python experiment.py
        ```

        If you already created a dataset from traces in the UI, replace the `create_dataset` call with:

        ```python theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        dataset = client.datasets.get_dataset(dataset="<YOUR DATASET NAME>")
        ```
      </Tab>

      <Tab title="TypeScript">
        ```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        npx tsx experiment.ts
        ```

        If you already created a dataset from traces in the UI, replace the `createDataset` call with:

        ```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
        import { getDataset } from "@arizeai/phoenix-client/datasets";

        const dataset = await getDataset({ client, dataset: { datasetName: "<YOUR DATASET NAME>" } });
        ```
      </Tab>
    </Tabs>
  </Step>

  <Step title={<span className="step-title">Review the scores</span>}>
    Head back to Phoenix and open the dataset's **Experiments** tab. You can compare outputs, evaluator labels, scores, and explanations for each example.

    <Frame caption="Side-by-side experiment comparison">
      <img src="https://storage.googleapis.com/arize-phoenix-assets/assets/images/get-started/evaluations-compare-score-delta.png" alt="Phoenix experiment comparison view showing a baseline support response scoring 0 and an edited response scoring 1" />
    </Frame>
  </Step>
</Steps>

## Mark a persistent baseline

The comparison above is transient. To pin one experiment as the reference point that later comparisons measure against, see [Set a Baseline Experiment](/docs/phoenix/datasets-and-experiments/how-to-experiments/running-experiments#set-a-baseline-experiment).
