Prerequisites
- A dataset in your space (Build a dataset).
- A registered agent configuration (Setting up your remote agent).
- Optional but recommended: tracing wired into your agent (Setting up tracing).
Launch an experiment
1
Open the dataset
Navigate to Datasets in the left nav and click the dataset you want to run against.
2
Click New Experiment → Run in Agent Playground
The agent playground modal opens with the dataset already selected.
3
Pick an agent
Use the Agent dropdown to select one of your registered agent configurations. The endpoint URL and auth from the configuration are used automatically — you don’t re-enter them per run.
4
Choose a preset (or write custom config)
If the agent has request presets, pick one from the Preset dropdown. The preset’s fields are merged into the request editor at the top level, so a preset that sets
config replaces config but leaves your goal placeholder untouched.You can edit the JSON before running — useful when you want to tweak one parameter from a known-good baseline without saving a new preset. Once you edit, Save as preset becomes available.5
Edit the request body
The Request editor holds the JSON Arize sends to your endpoint. It’s pre-filled from the agent’s input schema. Use These fields are sent at the top level of the request body, with
{{dataset.column_name}} to interpolate from each dataset row (the editor autocompletes column names):arize_metadata added beside them. The panel next to the editor previews the metadata Arize will inject.6
(Optional) Limit to a subset
For quick sanity checks, run on a subset of the dataset before committing to the full run. Useful for validating the request shape with one or two rows first.
7
(Optional) Add evaluators
Attach evaluators to score each run’s output automatically. See Evals overview for setup.
8
Click Run
Arize fans out the dataset against your endpoint in parallel, applies retries on transient failures, and streams results into a new experiment.
Watch a run in progress
As the experiment runs, the experiments view streams rows in real time:- Status column — pending, running, succeeded, failed.
- Output column — the agent’s response body.
- Latency / tokens — populated from the spans your agent emitted (if traced).
- Evaluator scores — computed as each row completes.
Inspecting an individual run
Click any row to see:- Input — exact JSON body sent to your endpoint: your hydrated fields plus the
arize_metadataArize appended. - Output — full response body returned.
- Trace — if your agent is traced, the linked trace tree (CHAIN, LLM, TOOL spans). Traces live in the space’s Agent Experiment Traces project, not your app’s tracing project.
- Headers — the request headers Arize sent, including
traceparent,Authorization, and any custom headers you configured. - Evaluator scores — per-evaluator pass/fail and reasoning.
Comparing runs
After you have two or more experiments on the same dataset, comparison is the point.From the experiments tab
Open Datasets → your dataset → Experiments, multi-select the runs you want to compare, and click Compare. The comparison view shows:- Side-by-side outputs for each dataset row across the selected runs.
- Evaluator deltas — which rows improved, regressed, or stayed flat.
- Summary metrics — pass rate, average latency, token counts per run.
- Tool-call patterns — if traced, you can see which runs called different tools or took different paths.
Common comparison patterns
Re-running failed rows
When a run finishes with failures, the Run button in the agent playground becomes Retry. Arize calls your endpoint again for the failed rows and merges the new results into the existing experiment — so you don’t have to lose the successful rows when chasing one flake.Rate limits, timeouts, and retries
These are set on the remote agent configuration, not per launch:- Rate limit (requests/minute) — caps how fast Arize calls your endpoint. Leave blank for the system default. Arize also backs off automatically when your endpoint returns
429. - Request timeout — per-request, in seconds. Default 120s, maximum 300s. Raise it for agents with long loops (e.g. multi-step research agents).
- Retries — transient failures (5xx, network errors) are retried automatically. This isn’t user-configurable.
Running from code or CLI
The agent playground is the UI path. To drive it from code (e.g. in CI), create aRUN_EXPERIMENT task with an AGENT_CALL run configuration. The run config takes the remote agent’s integration_id and an input_template (the same JSON you’d put in the request editor; {{column}} and {{dataset.column}} are equivalent).
- REST API —
POST /v2/taskswithrun_configuration.experiment_type: "AGENT_CALL". See the REST API reference. - Python SDK —
AgentCallRunConfigfromarize.tasks, passed asrun_configurationwhen creating a task. See Tasks (Python SDK). axCLI — register agents withax integrations create agent, then launch withax tasks create-run-experiment --run-configuration @run_config.json. Seeax tasks.
End-to-end example
Walking through the travel-agent demo (registered agent, dataset of 20 travel goals, three presets):1
Pick the dataset
Open
travel-goals-v1 (20 goals like “Plan a 3-day trip to Tokyo from SF in October”, “Weekend in NYC from Chicago”).2
Launch with the baseline preset
New Experiment → Run in Agent Playground → travel-agent → Production baseline (Sonnet 4.5) → Run. Wait ~3 minutes for 20 rows to complete.
3
Launch a second experiment with Opus
Same dataset, same agent, Opus 4.7 preset. Run. ~5 minutes (Opus is slower).
4
Compare
Compare Experiments → see per-row output diffs. Opus produced richer itineraries on 16/20 rows, but average latency was 2.3× higher. Pass rate on the “produces a coherent multi-day plan” evaluator: Sonnet 18/20, Opus 20/20.
5
Inspect a regression
On row 7 (Lisbon), Sonnet picked a $$$$ hotel; Opus picked a $$ one. Open the Sonnet trace, see the
search_hotels TOOL span — it ranked by rating, not by max_price constraint. Fix is a system prompt tweak.6
Iterate
Update the agent’s system prompt, redeploy, re-run the Sonnet experiment. Compare new Sonnet run to the previous one to confirm row 7 is fixed without breaking anything else.
Next
Compare experiments
Side-by-side diffs and evaluator deltas across runs.
Run evals on experiments
Add evaluators to score agent outputs.
CI/CD with experiments
Trigger agent experiments from your deploy pipeline.
Code experiments
Drive agent experiments from Python / TypeScript / CLI.