> ## Documentation Index
> Fetch the complete documentation index at: https://arize-ax.mintlify.site/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Importing data with GraphQL

> Automate file imports, table imports, integration keys, and BYOB connectors for S3, GCS, BigQuery, Snowflake, and Databricks with the Arize AX GraphQL API.

Arize AX's data import mutations let you create and manage the same ingestion jobs that the UI's [file](/docs/ax/machine-learning/machine-learning/how-to-ml/upload-data-to-arize/ui-drag-and-drop) and [table](/docs/ax/machine-learning/machine-learning/integrations-ml/google-bigquery) importers use, so you can provision them from a script instead of clicking through the UI every time you onboard a new model. This covers blob store file imports (S3, GCS, Azure), warehouse table imports (BigQuery, Snowflake, Databricks), the integration keys that power alert routing, and BYOB (bring your own bucket) connectors for Iceberg and Delta tables. For the connection setup each cloud provider requires before you can create a job, see the [S3](/docs/ax/machine-learning/machine-learning/integrations-ml/aws-s3-example), [GCS](/docs/ax/machine-learning/machine-learning/integrations-ml/gcs-example), [Azure](/docs/ax/machine-learning/machine-learning/integrations-ml/azure-example), [BigQuery](/docs/ax/machine-learning/machine-learning/integrations-ml/google-bigquery), [Snowflake](/docs/ax/machine-learning/machine-learning/integrations-ml/snowflake), and [Databricks](/docs/ax/machine-learning/machine-learning/integrations-ml/databricks) integration pages.

## Find the IDs you need

Every import job mutation takes a space ID, and `createIntegrationKey` takes an organization ID. Both come back from one query against `viewer`. See [using global node IDs](/docs/ax/graphql-reference/overview/how-to-use-graphql/using-global-node-ids) for how these opaque IDs work.

<CodeGroup>
  ```graphql Query theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  query FindSpaceAndOrg {
    viewer {
      spaces(first: 10, search: "Fraud") {
        edges { node { id name organization { id name } } }
      }
    }
  }
  ```
</CodeGroup>

Once you have a job's ID (from a create mutation response, or from listing jobs below), fetch it directly with `node(id: ...)`, selecting `... on FileImportJob` or `... on TableImportJob` for the job-specific fields.

## Create an integration key and test it

Integration keys connect Opsgenie, PagerDuty, or Slack so monitors can route alerts there. Create one under an organization, then call `testIntegrationKey` to confirm the provider accepts it before you wire it into a monitor. See [alerting integrations](/docs/ax/observe/production-monitoring/alerting-integrations) for the provider side of this setup.

<CodeGroup>
  ```graphql Create theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateIntegrationKey($input: CreateIntegrationKeyInput!) {
    createIntegrationKey(input: $input) {
      integrationKey { id name providerName alertSeverity }
    }
  }
  ```

  ```json Create variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "accountOrganizationId": "<ORGANIZATION_ID>",
      "providerName": "pagerduty",
      "serviceName": "ML Platform Team",
      "apiKey": "<PROVIDER_API_KEY>",
      "alertSeverity": "pagerdutycritical"
    }
  }
  ```

  ```graphql Test theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation TestIntegrationKey($integrationKeyId: ID!) {
    testIntegrationKey(input: { integrationKeyId: $integrationKeyId }) {
      statusCode
    }
  }
  ```

  ```json Test variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "integrationKeyId": "<INTEGRATION_KEY_ID>" }
  ```
</CodeGroup>

A `statusCode` of `200` or `201` means the provider accepted the key. Anything in the 400s or 401 means the API key or service name is wrong on the provider's side.

Reference: [`createIntegrationKey`](/docs/ax/graphql-reference/mutations/data-import#createintegrationkey), [`testIntegrationKey`](/docs/ax/graphql-reference/mutations/data-import#testintegrationkey).

## Create a file import job from cloud storage

`createFileImportJob` points Arize at a bucket and prefix, and the `schema` input maps your column names to model dimensions. Set `dryRun: true` first to validate the mapping with no data written; the `validationResult` field reports back per-file errors. Switch `dryRun` to `false` (or omit it, since it defaults to `false`) once the dry run is clean.

<CodeGroup>
  ```graphql Mutation theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateFileImportJob($input: CreateFileImportJobInput!) {
    createFileImportJob(input: $input) {
      fileImportJob { id jobId jobStatus }
      validationResult {
        validationStatus
        filePath
        error { code message row }
      }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "spaceId": "<SPACE_ID>",
      "modelName": "fraud-detection-model",
      "modelType": "score_categorical",
      "modelEnvironmentName": "production",
      "blobStore": "S3",
      "bucketName": "my-arize-exports",
      "prefix": "fraud-model/production/",
      "schema": { "predictionId": "prediction_id", "predictionLabel": "prediction_label", "actualLabel": "actual_label", "timestamp": "prediction_ts", "featuresList": ["amount", "merchant_category", "account_age_days"] },
      "dryRun": true
    }
  }
  ```
</CodeGroup>

For Azure buckets accessed through a service principal or managed identity, also pass `azureStorageIdentifier: { tenantId, storageAccountName }`.

Reference: [`createFileImportJob`](/docs/ax/graphql-reference/mutations/data-import#createfileimportjob).

## Manage a file import job's lifecycle

Once a job exists, you poll it for progress, retry individual failed files, and pause or resume it without recreating it.

Query the job through `node`. `files` accepts a `status` filter (`PENDING`, `COMPLETE`, `FAILED`, `SKIPPED`, `CANCELED`) and an optional `startTime`/`endTime` window.

<CodeGroup>
  ```graphql Query theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  query PollFileImportJob($jobId: ID!) {
    node(id: $jobId) {
      ... on FileImportJob {
        jobStatus
        totalFilesPendingCount
        totalFilesFailedCount
        files(first: 25, status: FAILED) {
          edges { node { id filePath error { message } } }
        }
      }
    }
  }
  ```
</CodeGroup>

The same call works from a script, so you can poll in a loop until nothing is pending:

```python theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
import time
import requests

QUERY = """
query PollFileImportJob($jobId: ID!) {
  node(id: $jobId) { ... on FileImportJob { jobStatus totalFilesPendingCount } }
}
"""

def poll(job_id):
    while True:
        resp = requests.post("https://app.arize.com/graphql",
            headers={"x-api-key": "<YOUR_API_KEY>"},
            json={"query": QUERY, "variables": {"jobId": job_id}})
        node = resp.json()["data"]["node"]
        if node["totalFilesPendingCount"] == 0:
            return node
        time.sleep(30)
```

`resetFileStatus` takes the job ID and a file ID (from the `files` query above) and queues that file for another attempt. `pauseFileImportJob` and `startFileImportJob` (resume) take just the job ID.

<CodeGroup>
  ```graphql Reset theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation ResetFileStatus($input: ResetFileStatusInput!) {
    resetFileStatus(input: $input) {
      job { jobStatus totalFilesPendingCount }
    }
  }
  ```

  ```json Reset variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "input": { "jobId": "<JOB_ID>", "fileId": "<FILE_ID>" } }
  ```

  ```graphql Pause theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation PauseFileImportJob($jobId: ID!) {
    pauseFileImportJob(input: { jobId: $jobId }) {
      fileImportJob { jobStatus }
    }
  }
  ```

  ```json Pause variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "jobId": "<JOB_ID>" }
  ```
</CodeGroup>

Resume with `startFileImportJob(input: { jobId: $jobId })`, the same input shape as `pauseFileImportJob`.

Reference: [`resetFileStatus`](/docs/ax/graphql-reference/mutations/data-import#resetfilestatus), [`pauseFileImportJob`](/docs/ax/graphql-reference/mutations/data-import#pausefileimportjob), [`startFileImportJob`](/docs/ax/graphql-reference/mutations/data-import#startfileimportjob).

## Create a table import job

`createTableImportJob` works the same way as the file import job, except you supply one of `bigQueryTableConfig`, `snowflakeTableConfig`, or `databricksTableConfig` alongside `tableStore` instead of a bucket and prefix.

<CodeGroup>
  ```graphql Mutation theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateTableImportJob($input: CreateTableImportJobInput!) {
    createTableImportJob(input: $input) {
      tableImportJob { id jobId jobStatus }
      validationResult { validationStatus error { message } }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "spaceId": "<SPACE_ID>",
      "modelName": "credit-risk-model",
      "modelType": "score_categorical",
      "modelEnvironmentName": "production",
      "tableStore": "BigQuery",
      "bigQueryTableConfig": { "projectId": "my-gcp-project", "dataset": "arize_exports", "tableName": "credit_risk_predictions" },
      "schema": { "predictionId": "prediction_id", "predictionLabel": "prediction_label", "timestamp": "prediction_ts", "features": "feature_", "actualLabel": "actual_label", "changeTimestamp": "partition_timestamp" },
      "dryRun": true
    }
  }
  ```
</CodeGroup>

For Snowflake, replace `bigQueryTableConfig` with `snowflakeTableConfig: { accountID, schema, database, tableName }`. For Databricks, use `databricksTableConfig: { hostname, endpoint, port, catalog, databricksSchema, tableName, token }` (or `azureResourceId` instead of `token` for Azure Databricks workspaces).

Reference: [`createTableImportJob`](/docs/ax/graphql-reference/mutations/data-import#createtableimportjob).

## Update ingestion parameters for a table import job

`updateTableIngestionParameters` changes how often an existing job queries the warehouse without touching its schema. Note the unit difference described in the Gotchas section below: the input takes minutes and hours, but reading the job back reports seconds.

<CodeGroup>
  ```graphql Mutation theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation UpdateTableIngestionParameters($input: UpdateTableIngestionParametersInput!) {
    updateTableIngestionParameters(input: $input) {
      tableImportJob {
        id
        tableIngestionParameters { refreshIntervalSeconds queryWindowSizeSeconds }
      }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "jobId": "<JOB_ID>",
      "tableIngestionParameters": { "refreshIntervalMinutes": 60, "queryWindowSizeHours": 24 }
    }
  }
  ```
</CodeGroup>

Reference: [`updateTableIngestionParameters`](/docs/ax/graphql-reference/mutations/data-import#updatetableingestionparameters).

## Schedule a triggered run for an ongoing Snowflake table job

Event-based table ingestion is available for Snowflake only. Create the job with `createTriggeredOngoingTableImportJob`, look up its `jobId` from `tableJobs` on the space (see the next recipe), then trigger each run for a specific time window with `createTriggeredOngoingTableRun`.

<CodeGroup>
  ```graphql Create job theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateTriggeredOngoingTableJob($input: CreateTriggeredOngoingTableImportJobInput!) {
    createTriggeredOngoingTableImportJob(input: $input) {
      tableImportJob { jobId jobStatus }
      validationResult { validationStatus error { message } }
    }
  }
  ```

  ```json Create job variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "spaceId": "<SPACE_ID>",
      "modelName": "inventory-forecast-model",
      "modelType": "score_categorical",
      "modelEnvironmentName": "production",
      "tableStore": "Snowflake",
      "snowflakeTableConfig": { "accountID": "my-snowflake-account", "schema": "public", "database": "analytics", "tableName": "inventory_predictions" },
      "schema": { "predictionId": "prediction_id", "predictionLabel": "prediction_label", "timestamp": "prediction_ts", "featuresList": ["sku", "warehouse_region"], "changeTimestamp": "partition_timestamp" }
    }
  }
  ```

  ```graphql Trigger run theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation TriggerTableRun($input: CreateTriggeredOngoingTableRunInput!) {
    createTriggeredOngoingTableRun(input: $input) {
      clientMutationId
    }
  }
  ```

  ```json Trigger run variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "input": { "jobId": "<JOB_ID>", "queryStart": "2026-09-20T01:00:00Z", "queryEnd": "2026-09-20T23:00:00Z" } }
  ```
</CodeGroup>

Reference: [`createTriggeredOngoingTableImportJob`](/docs/ax/graphql-reference/mutations/data-import#createtriggeredongoingtableimportjob), [`createTriggeredOngoingTableRun`](/docs/ax/graphql-reference/mutations/data-import#createtriggeredongoingtablerun).

## List and delete import jobs

List jobs off the `Space` node, then delete the ones you no longer need by `jobId`. This works the same way for file jobs (`importJobs`) and table jobs (`tableJobs`).

<CodeGroup>
  ```graphql List theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  query ListTableJobs($spaceId: ID!) {
    node(id: $spaceId) {
      ... on Space {
        tableJobs(first: 50) {
          edges { node { id modelName jobStatus totalQueriesFailedCount totalQueriesSuccessCount } }
        }
      }
    }
  }
  ```

  ```graphql Delete theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation DeleteTableImportJob($jobId: ID!) {
    deleteTableImportJob(input: { jobId: $jobId }) {
      tableImportJob { jobStatus }
    }
  }
  ```

  ```json Delete variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "jobId": "<JOB_ID>" }
  ```
</CodeGroup>

Use `deleteFileImportJob(input: { jobId })` the same way for file import jobs.

Reference: [`deleteTableImportJob`](/docs/ax/graphql-reference/mutations/data-import#deletetableimportjob), [`deleteFileImportJob`](/docs/ax/graphql-reference/mutations/data-import#deletefileimportjob).

## Validate and create a BYOB connector

BYOB (bring your own bucket) connectors read Iceberg or Delta tables directly out of a bucket you own. Validate the bucket and path before creating the connector, since a bad path or missing bucket permissions is the most common failure.

<CodeGroup>
  ```graphql Validate theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation ValidateByobConnector($input: ValidateByobConnectorInput!) {
    validateByobConnector(input: $input) {
      isValid
      validationReason
      validationMessage
    }
  }
  ```

  ```json Validate variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "spaceId": "<SPACE_ID>",
      "connectorIdentifier": { "spaceId": "<SPACE_ID>", "cloudStorageProvider": "S3", "bucketAndPath": "my-iceberg-bucket/warehouse/orders" }
    }
  }
  ```

  ```graphql Create theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateByobConnector($input: CreateByobConnectorInput!) {
    createByobConnector(input: $input) {
      id
    }
  }
  ```

  ```json Create variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "spaceId": "<SPACE_ID>",
      "connectorIdentifier": { "spaceId": "<SPACE_ID>", "cloudStorageProvider": "S3", "bucketAndPath": "my-iceberg-bucket/warehouse/orders" },
      "name": "orders-iceberg-connector",
      "externalModelIds": ["orders-fraud-model"],
      "tableFormat": "ICEBERG"
    }
  }
  ```
</CodeGroup>

Reference: [`validateByobConnector`](/docs/ax/graphql-reference/mutations/data-import#validatebyobconnector), [`createByobConnector`](/docs/ax/graphql-reference/mutations/data-import#createbyobconnector).

## Gotchas and behavior notes

<AccordionGroup>
  <Accordion title="The job status enum you set is not the one you read back">
    `updateFileImportJob` and `updateTableImportJob` take a `jobStatus` of type `JobStatus`, which only has two values: `ONGOING_INACTIVE` and `ONGOING_DELETED`. But when you read `jobStatus` back off a `FileImportJob` or `TableImportJob`, it's typed `ImportJobStatus`, with values `active`, `inactive`, and `deleted`. These are two different enums for the same concept. Don't reuse the value you read back as the value you write.
  </Accordion>

  <Accordion title="Table ingestion parameters change units between input and output">
    `TableIngestionParametersInputType` (used by `createTableImportJob`, `updateTableImportJob`, and `updateTableIngestionParameters`) takes `refreshIntervalMinutes` and `queryWindowSizeHours`. Reading the same settings back through `TableImportJob.tableIngestionParameters` returns `TableIngestionParametersType`, with `refreshIntervalSeconds` and `queryWindowSizeSeconds` instead. Convert accordingly if you round-trip a value.
  </Accordion>

  <Accordion title="Update mutations require the full schema, not a partial patch">
    `UpdateFileImportJobInput.schema` and `UpdateTableImportJobInput.schema` are both non-null, and `dryRun` defaults to `false` on the create mutations (set it to `true` to validate with nothing written). So if you only want to change `jobStatus` or `modelVersion` on an existing job, you still have to resend the complete `schema` input or the mutation is rejected.
  </Accordion>

  <Accordion title="BYOB has two soft spots: nullable deletes and a silent default">
    `DeleteByobConnectorPayload.success` and `DeleteByobDatasourcePayload.success` are nullable `Boolean`, not `Boolean!`, so treat a missing value as a failure and re-check with a query. Separately, `CreateByobConnectorInput.tableFormat` is optional and defaults to `ICEBERG` per its field description; pass `DELTA` explicitly for Delta Lake tables.
  </Accordion>
</AccordionGroup>

<CardGroup cols={3}>
  <Card title="Data import mutations" icon="book" href="/docs/ax/graphql-reference/mutations/data-import">
    Full argument and return type reference for every mutation used above.
  </Card>

  <Card title="All mutations" icon="list" href="/docs/ax/graphql-reference/mutations">
    Index of every mutation domain in the Arize GraphQL API.
  </Card>

  <Card title="API Explorer" icon="terminal" href="/docs/ax/graphql-reference/overview/api-explorer">
    Run these queries and mutations interactively against your own space.
  </Card>
</CardGroup>
