Run an experiment
1
Load a prompt and dataset
Open Playground, load your prompt, then select a dataset (or replay a production span).
2
Attach evaluators
Add evaluators so each output is scored automatically.
3
Run and inspect
Click Run, then open View Experiment to review outputs, latency, tokens, and evaluator scores.
Compare experiments
Once you have multiple runs on the same dataset, open Compare Experiments to inspect:- Output differences
- Evaluator deltas
- Summary metrics by run
- Regressions vs baseline
Playground views
A Playground View is a named snapshot of your Playground session. It stores:- Prompt messages and tools
- Model and parameter configuration
- Dataset or span context
- Generated results and scores

Save and manage Playground views from the Playgrounds list