Skip to main content
Arize AX’s data import mutations let you create and manage the same ingestion jobs that the UI’s file and table importers use, so you can provision them from a script instead of clicking through the UI every time you onboard a new model. This covers blob store file imports (S3, GCS, Azure), warehouse table imports (BigQuery, Snowflake, Databricks), the integration keys that power alert routing, and BYOB (bring your own bucket) connectors for Iceberg and Delta tables. For the connection setup each cloud provider requires before you can create a job, see the S3, GCS, Azure, BigQuery, Snowflake, and Databricks integration pages.

Find the IDs you need

Every import job mutation takes a space ID, and createIntegrationKey takes an organization ID. Both come back from one query against viewer. See using global node IDs for how these opaque IDs work.
Once you have a job’s ID (from a create mutation response, or from listing jobs below), fetch it directly with node(id: ...), selecting ... on FileImportJob or ... on TableImportJob for the job-specific fields.

Create an integration key and test it

Integration keys connect Opsgenie, PagerDuty, or Slack so monitors can route alerts there. Create one under an organization, then call testIntegrationKey to confirm the provider accepts it before you wire it into a monitor. See alerting integrations for the provider side of this setup.
A statusCode of 200 or 201 means the provider accepted the key. Anything in the 400s or 401 means the API key or service name is wrong on the provider’s side. Reference: createIntegrationKey, testIntegrationKey.

Create a file import job from cloud storage

createFileImportJob points Arize at a bucket and prefix, and the schema input maps your column names to model dimensions. Set dryRun: true first to validate the mapping with no data written; the validationResult field reports back per-file errors. Switch dryRun to false (or omit it, since it defaults to false) once the dry run is clean.
For Azure buckets accessed through a service principal or managed identity, also pass azureStorageIdentifier: { tenantId, storageAccountName }. Reference: createFileImportJob.

Manage a file import job’s lifecycle

Once a job exists, you poll it for progress, retry individual failed files, and pause or resume it without recreating it. Query the job through node. files accepts a status filter (PENDING, COMPLETE, FAILED, SKIPPED, CANCELED) and an optional startTime/endTime window.
The same call works from a script, so you can poll in a loop until nothing is pending:
resetFileStatus takes the job ID and a file ID (from the files query above) and queues that file for another attempt. pauseFileImportJob and startFileImportJob (resume) take just the job ID.
Resume with startFileImportJob(input: { jobId: $jobId }), the same input shape as pauseFileImportJob. Reference: resetFileStatus, pauseFileImportJob, startFileImportJob.

Create a table import job

createTableImportJob works the same way as the file import job, except you supply one of bigQueryTableConfig, snowflakeTableConfig, or databricksTableConfig alongside tableStore instead of a bucket and prefix.
For Snowflake, replace bigQueryTableConfig with snowflakeTableConfig: { accountID, schema, database, tableName }. For Databricks, use databricksTableConfig: { hostname, endpoint, port, catalog, databricksSchema, tableName, token } (or azureResourceId instead of token for Azure Databricks workspaces). Reference: createTableImportJob.

Update ingestion parameters for a table import job

updateTableIngestionParameters changes how often an existing job queries the warehouse without touching its schema. Note the unit difference described in the Gotchas section below: the input takes minutes and hours, but reading the job back reports seconds.
Reference: updateTableIngestionParameters.

Schedule a triggered run for an ongoing Snowflake table job

Event-based table ingestion is available for Snowflake only. Create the job with createTriggeredOngoingTableImportJob, look up its jobId from tableJobs on the space (see the next recipe), then trigger each run for a specific time window with createTriggeredOngoingTableRun.
Reference: createTriggeredOngoingTableImportJob, createTriggeredOngoingTableRun.

List and delete import jobs

List jobs off the Space node, then delete the ones you no longer need by jobId. This works the same way for file jobs (importJobs) and table jobs (tableJobs).
Use deleteFileImportJob(input: { jobId }) the same way for file import jobs. Reference: deleteTableImportJob, deleteFileImportJob.

Validate and create a BYOB connector

BYOB (bring your own bucket) connectors read Iceberg or Delta tables directly out of a bucket you own. Validate the bucket and path before creating the connector, since a bad path or missing bucket permissions is the most common failure.
Reference: validateByobConnector, createByobConnector.

Gotchas and behavior notes

updateFileImportJob and updateTableImportJob take a jobStatus of type JobStatus, which only has two values: ONGOING_INACTIVE and ONGOING_DELETED. But when you read jobStatus back off a FileImportJob or TableImportJob, it’s typed ImportJobStatus, with values active, inactive, and deleted. These are two different enums for the same concept. Don’t reuse the value you read back as the value you write.
TableIngestionParametersInputType (used by createTableImportJob, updateTableImportJob, and updateTableIngestionParameters) takes refreshIntervalMinutes and queryWindowSizeHours. Reading the same settings back through TableImportJob.tableIngestionParameters returns TableIngestionParametersType, with refreshIntervalSeconds and queryWindowSizeSeconds instead. Convert accordingly if you round-trip a value.
UpdateFileImportJobInput.schema and UpdateTableImportJobInput.schema are both non-null, and dryRun defaults to false on the create mutations (set it to true to validate with nothing written). So if you only want to change jobStatus or modelVersion on an existing job, you still have to resend the complete schema input or the mutation is rejected.
DeleteByobConnectorPayload.success and DeleteByobDatasourcePayload.success are nullable Boolean, not Boolean!, so treat a missing value as a failure and re-check with a query. Separately, CreateByobConnectorInput.tableFormat is optional and defaults to ICEBERG per its field description; pass DELTA explicitly for Delta Lake tables.

Data import mutations

Full argument and return type reference for every mutation used above.

All mutations

Index of every mutation domain in the Arize GraphQL API.

API Explorer

Run these queries and mutations interactively against your own space.