Skip to main content
Datasets store versioned example sets you run experiments against, and annotations let your team attach human labels and scores to spans, traces, sessions, dataset records and experiment runs. This guide covers the GraphQL mutations for tagging datasets and experiments, editing dataset version examples, and creating and populating annotation configs and queues so you can automate the human review workflow instead of clicking through the UI.

Find the IDs you need

Every mutation on this page takes a space-scoped ID (a dataset, experiment, dataset version or model/project ID). Start from viewer to list the spaces your API key can see, then pull the datasets, experiments and tags in a space. See using global node IDs for how these opaque IDs work.
To get the IDs of individual examples inside a dataset version (needed for patches and deletes), expand latestDatasetVersion { id examples(first: 50) { edges { node { id } } } } on a Dataset.

Tag a dataset or experiment

Tags group datasets and experiments so you can filter for them later (for example, in the tagsAny or tagsAll arguments on Space.datasets and Space.experiments). Create the tags first with Space.tags, or reuse existing tag IDs, then attach them.
Tagging an experiment uses the same shape with addTagsToExperiment and experimentId. Both mutations return a TagAssociationSuccess or TagAssociationError union, so check __typename rather than assuming success. Remove tags the same way with removeTagsFromDataset or removeTagsFromExperiment. Reference: addTagsToDataset, addTagsToExperiment.

Add production spans to a dataset version

Copy records straight from a project’s traces into a dataset version without exporting and re-uploading anything. This mirrors the curate a dataset workflow but lets you script the selection with a time range and filters.
An empty recordIds list under records selects every record in the interval that matches the set’s filters, so add a queryFilter (for example "annotation.Correctness.label == 'incorrect'") when you only want a slice of traffic. Use sessions.sessionIds instead of records.recordIds to copy whole sessions, and deletedRecordIds to remove examples. recordIds at the top level of NewExampleSet is deprecated in favor of records. Reference: updateDatasetVersionExamples.

Patch an example’s columns

Fix a value in an existing dataset example (for example, correcting an expected output) without re-adding the record.
value is always a string. dataType tells the server how to parse it, so a LONG or FLOAT column still takes a quoted numeric string. Reference: updateDatasetVersionExamples.

Create an annotation config

Annotation configs define the label schema annotators fill in. Pick categorical for a fixed label set, continuous for a numeric score range, or freeform for unstructured text notes with no extra config block.
For a continuous config, swap categoricalConfig for continuousConfig: { minValue: 1, maxValue: 5, optimizationDirection: maximize }. For freeform, set annotationConfigType: "freeform" and omit both config blocks. Reference: createAnnotationConfig.

Create an annotation queue for human review

Queues route a set of records to one or more annotators. You can create new annotation configs inline, reuse existing ones by ID, or both, and populate the queue directly from spans in a project’s trace data.
assignmentMethod: all gives every assigned annotator every record; random spreads records across them. The older datasetVersionId and top-level recordIds inputs are deprecated in favor of recordSpecifications, which also supports pulling straight from a dataset version with sourceType: "dataset". Use columnAllowlist to limit which columns annotators see. Reference: createAnnotationQueue.

Add annotations to spans in bulk

Once your subject-matter experts have labeled data in an external tool (a spreadsheet, an internal review app), write the results back onto the matching spans in one call. batchUpdateAnnotations takes a single model/project ID and a list of per-record updates, each with its own annotation config and label or score.
startTime is a search-space filter, not just metadata: it must be at most 24 hours before the record’s actual timestamp or the record will not be found. annotationType is Label, Score or Text (capitalized), and only the matching field (label, score or text) on annotation is read. For a single record, use updateAnnotations with a modelRecordContext, experimentRunContext or datasetRecordContext instead of a model-wide batch. Reference: batchUpdateAnnotations, updateAnnotations.

Python: bulk-import annotations from a spreadsheet

The most common scripted use of this API is exporting labels from an external review tool and pushing them into Arize AX. This sends one batchUpdateAnnotations call per model.

Gotchas and behavior notes

addTagsToDataset, removeTagsFromDataset, addTagsToExperiment, removeTagsFromExperiment, updateAnnotations and batchUpdateAnnotations all return a success/error union in their result field instead of throwing a GraphQL error on a logical failure. Always select both branches with inline fragments, as shown above, and check __typename.
The AnnotationType enum is Label, Score, Text, not the lowercase label/score used in some older examples. Only populate the field on AnnotationInput that matches the chosen type.
The note input on updateAnnotations, batchUpdateAnnotations (per record) and the annotation queue note path is deprecated in favor of a freeform-text annotation. Create a freeform annotation config and write the text through AnnotationInput.text with annotationType: "Text" instead.
NewExampleSet.recordIds (on updateDatasetVersionExamples) is superseded by records.recordIds, and CreateAnnotationQueueInput.recordIds / datasetVersionId are superseded by recordSpecifications. Both older fields still work but new integrations should use the replacement.
On ModelRecordContextInput and RecordAnnotationUpdateInput, startTime narrows the server-side search for the record and must be close to (at most 24 hours before) the record’s real timestamp, or the mutation will not find the record to annotate.
CreateAnnotationQueueInput.maxRecords must be between 1 and 5000.

Dataset and annotation mutations

Full argument and return-type reference for all nine mutations.

All mutations

Index of every GraphQL mutation grouped by domain.

API explorer

Run queries and mutations interactively against your own account.