> ## Documentation Index
> Fetch the complete documentation index at: https://arize-ax.mintlify.site/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Managing evaluators and tasks with GraphQL

> Automate Arize AX evals with the GraphQL API: create judges in Eval Hub, bind them to online tasks with sampling rates, and run evaluations on demand.

An [evaluator](/docs/ax/evaluate/create-evaluators) defines what to measure (an LLM-as-judge template, code, or an agent-as-a-judge harness). A [task](/docs/ax/evaluate/run-evals) points that evaluator at data: a project's live traces, a dataset, or an experiment, and controls how much of it gets scored and how often. The GraphQL API covers both halves, plus the related task types for running an experiment on demand and driving prompt optimization.

This guide assumes you already know how to [form a GraphQL call](/docs/ax/graphql-reference/overview/how-to-use-graphql/forming-calls) with an `x-api-key` header against `https://app.arize.com/graphql`.

## Find the IDs you need

Evaluator and task mutations need a space ID, a model (project) ID, and often an evaluator ID or a named LLM integration ID. Start from `viewer`.

<CodeGroup>
  ```graphql Query theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  query FindResourcesForTasks($search: String) {
    viewer {
      spaces(search: $search, first: 5) {
        edges {
          node {
            id
            name
            models(first: 10) {
              edges { node { id name } }
            }
            evaluators(first: 10) {
              edges { node { id name taskType } }
            }
            llmIntegrations {
              id
              provider
            }
          }
        }
      }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "search": "LLM_test"
  }
  ```
</CodeGroup>

Once you have an evaluator, task, or run ID, fetch it directly with `node(id: "<ID>") { ... on Evaluator { name } }`. See [Using global node IDs](/docs/ax/graphql-reference/overview/how-to-use-graphql/using-global-node-ids) for how these opaque IDs work.

## List evaluators and online tasks in a space

Evaluators live in Eval Hub independently of any task; a single evaluator can be attached to many tasks. This query lists both so you can see what exists before you wire one to the other.

<CodeGroup>
  ```graphql Query theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  query ListEvaluatorsAndTasks($spaceId: ID!) {
    node(id: $spaceId) {
      ... on Space {
        evaluators(first: 20) {
          edges { node { id name taskType commitMessage } }
        }
        onlineTasks(first: 20) {
          edges { node { id name taskType samplingRate runContinuously lastRunAt } }
        }
      }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "spaceId": "<SPACE_ID>"
  }
  ```
</CodeGroup>

## Create an LLM-as-judge evaluator and a new version

`createEvaluator` creates the evaluator and its first version together. `createEvaluatorVersion` adds a new version (for example, tightening the rails) without losing history. To just rename or re-describe an evaluator without creating a version, use [`editEvaluator`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#editevaluator) instead; it takes `evaluatorId`, `name`, and `description`.

<CodeGroup>
  ```graphql Create theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateEvaluator($input: CreateEvaluatorMutationInput!) {
    createEvaluator(input: $input) {
      evaluator { id name }
    }
  }
  ```

  ```graphql New version theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateEvaluatorVersion($input: CreateEvaluatorVersionMutationInput!) {
    createEvaluatorVersion(input: $input) {
      evaluatorVersion { id versionNumber }
    }
  }
  ```

  ```json Create variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "spaceId": "<SPACE_ID>",
      "name": "Answer correctness",
      "description": "Checks whether the response correctly answers the question.",
      "commitMessage": "Initial version",
      "templateEvaluator": {
        "name": "correctness",
        "rails": ["correct", "incorrect"],
        "template": "Given the question {question} and the response {response}, is the response correct? Answer correct or incorrect.",
        "position": 0,
        "includeExplanations": true,
        "useFunctionCallingIfAvailable": false
      }
    }
  }
  ```

  ```json New version variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "evaluatorId": "<EVALUATOR_ID>",
      "commitMessage": "Add a partially_correct rail",
      "templateEvaluator": {
        "name": "correctness",
        "rails": ["correct", "partially_correct", "incorrect"],
        "template": "Given the question {question} and the response {response}, is the response correct, partially_correct, or incorrect?",
        "position": 0,
        "includeExplanations": true,
        "useFunctionCallingIfAvailable": false
      }
    }
  }
  ```
</CodeGroup>

Reference: [`createEvaluator`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#createevaluator), [`createEvaluatorVersion`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#createevaluatorversion).

## Create an online eval task bound to an evaluator

`createEvalTask` points an evaluator at live traces from a project. `samplingRate` is a fraction between 0 and 1, and `queryFilter` narrows which spans are admitted before sampling is applied.

<CodeGroup>
  ```graphql Mutation theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreateEvalTask($input: CreateEvalTaskMutationInput!) {
    createEvalTask(input: $input) {
      evalTask { id name samplingRate runContinuously }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "modelId": "<MODEL_ID>",
      "name": "Correctness on production traces",
      "samplingRate": 0.2,
      "runContinuously": true,
      "queryFilter": "attributes.openinference.span.kind = 'LLM'",
      "evaluators": [{ "evaluatorId": "<EVALUATOR_ID>", "position": 0 }]
    }
  }
  ```
</CodeGroup>

Reference: [`createEvalTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#createevaltask).

## Pause, resume, resample, or migrate a task to Eval Hub

`patchEvalTask` updates any subset of a task's fields; omit what you are not changing. Older tasks created before Eval Hub store their evaluator config inline on the task; [`migrateTaskToEvalHub`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#migratetasktoevalhub) (input: just `onlineTaskId`) converts one to reference Eval Hub evaluators instead. Once migrated, use `patchEvalTaskWithEvalHub` rather than `patchEvalTask`, since it takes `evaluators` (hub references) instead of inline `templateEvaluators`/`codeEvaluators`.

<CodeGroup>
  ```graphql Pause theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation PatchEvalTask($input: PatchEvalTaskMutationInput!) {
    patchEvalTask(input: $input) {
      evalTask { id name samplingRate runContinuously }
    }
  }
  ```

  ```graphql Patch with Eval Hub theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation PatchEvalTaskWithEvalHub($input: PatchEvalTaskWithEvalHubMutationInput!) {
    patchEvalTaskWithEvalHub(input: $input) {
      evalTask { id name samplingRate }
    }
  }
  ```

  ```json Pause variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "input": { "onlineTaskId": "<TASK_ID>", "runContinuously": false } }
  ```

  ```json Patch with Eval Hub variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "onlineTaskId": "<TASK_ID>",
      "runContinuously": true,
      "samplingRate": 0.5,
      "evaluators": [{ "evaluatorId": "<EVALUATOR_ID>", "position": 0 }]
    }
  }
  ```
</CodeGroup>

Reference: [`patchEvalTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#patchevaltask), [`patchEvalTaskWithEvalHub`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#patchevaltaskwithevalhub).

## Run a task on demand and cancel a run

`runOnlineTask` backfills a task over a historical time range instead of waiting for live traffic, capped at `maxSpans` (10,000 by default). Both mutations return a union: a success type on the happy path, or `TaskError` with a `message` and `code`.

<CodeGroup>
  ```graphql Run theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation RunOnlineTask($input: RunOnlineTaskMutationInput!) {
    runOnlineTask(input: $input) {
      result {
        __typename
        ... on CreateTaskRunResponse {
          taskRun { ... on EvaluationTaskRun { id status } }
        }
        ... on TaskError { message code }
      }
    }
  }
  ```

  ```graphql Cancel theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CancelOnlineTaskRun($input: CancelOnlineTaskRunMutationInput!) {
    cancelOnlineTaskRun(input: $input) {
      result {
        __typename
        ... on CancelTaskRunResponse {
          taskRun { ... on EvaluationTaskRun { id status } }
        }
        ... on TaskError { message code }
      }
    }
  }
  ```

  ```json Run variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "onlineTaskId": "<TASK_ID>",
      "dataStartTime": "2026-01-01T00:00:00Z",
      "dataEndTime": "2026-01-02T00:00:00Z",
      "maxSpans": 500
    }
  }
  ```

  ```json Cancel variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "input": { "onlineTaskRunId": "<RUN_ID>" } }
  ```
</CodeGroup>

Reference: [`runOnlineTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#runonlinetask), [`cancelOnlineTaskRun`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#cancelonlinetaskrun).

## Run an experiment task

`runExperimentTask` runs a configured LLM generation (or a template eval, code eval, or agent call) across a dataset or dataset version, one `runConfigurations` entry per instance you want to compare. Only the required fields are shown; add `templateConfig`, `codeEvalConfig`, or `agentConfig` to a run configuration depending on its `experimentType`.

<CodeGroup>
  ```graphql Mutation theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation RunExperimentTask($input: RunExperimentTaskMutationInput!) {
    runExperimentTask(input: $input) {
      result {
        __typename
        ... on RunExperimentTaskSuccess {
          taskRun { ... on RunExperimentTaskRun { id } }
        }
        ... on RunExperimentTaskError { message code }
      }
    }
  }
  ```

  ```json Variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "datasetId": "<DATASET_ID>",
      "datasetName": "support-eval-set",
      "datasetVersionId": "<DATASET_VERSION_ID>",
      "spaceId": "<SPACE_ID>",
      "exampleOverrides": [
        { "exampleId": "<EXAMPLE_ID>", "overridesJSON": {} }
      ],
      "runConfigurations": [
        {
          "experimentType": "llm_generation",
          "experimentName": "gpt-5.5 baseline",
          "instanceId": 1
        }
      ]
    }
  }
  ```
</CodeGroup>

Reference: [`runExperimentTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#runexperimenttask).

## Create and update a prompt optimization task

`createPromptOptimizationTask` runs [prompt learning](/docs/ax/prompts/prompt-optimization) against a dataset (or an experiment's output), using feedback columns to steer the rewrite. `patchPromptOptimizationTask` updates it, most commonly to turn continuous optimization on or off.

<CodeGroup>
  ```graphql Create theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation CreatePromptOptimizationTask($input: CreatePromptOptimizationTaskMutationInput!) {
    createPromptOptimizationTask(input: $input) {
      promptOptimizationTask { id name }
    }
  }
  ```

  ```graphql Patch theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation PatchPromptOptimizationTask($input: PatchPromptOptimizationTaskMutationInput!) {
    patchPromptOptimizationTask(input: $input) {
      promptOptimizationTask { id runContinuously }
    }
  }
  ```

  ```json Create variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "datasetId": "<DATASET_ID>",
      "name": "Optimize support-reply-drafter",
      "runContinuously": false,
      "llmConfig": {
        "integrationId": "<INTEGRATION_ID>",
        "modelName": "gpt-4.1",
        "invocationParameters": { "temperature": 0 },
        "providerParameters": {}
      },
      "promptOptimizationConfig": {
        "userInstructions": "Make replies more concise and always propose a next step.",
        "outputColumn": "response",
        "feedbackColumns": ["correctness", "tone"],
        "batchSize": 20,
        "promptId": "<PROMPT_ID>"
      }
    }
  }
  ```

  ```json Patch variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  {
    "input": {
      "onlineTaskId": "<TASK_ID>",
      "runContinuously": true
    }
  }
  ```
</CodeGroup>

Reference: [`createPromptOptimizationTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#createpromptoptimizationtask), [`patchPromptOptimizationTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#patchpromptoptimizationtask).

## Delete an evaluator and online tasks

Deleting an evaluator does not delete the tasks it was attached to; detach or delete those tasks separately. `deleteOnlineTask` is a batch mutation, up to 50 task IDs per request.

<CodeGroup>
  ```graphql Delete evaluator theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation DeleteEvaluator($input: DeleteEvaluatorMutationInput!) {
    deleteEvaluator(input: $input) {
      success
    }
  }
  ```

  ```graphql Delete tasks theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  mutation DeleteOnlineTask($input: DeleteOnlineTaskMutationInput!) {
    deleteOnlineTask(input: $input) {
      success
      deletedTaskIds
      failedTasks { taskId message }
    }
  }
  ```

  ```json Delete evaluator variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "input": { "evaluatorId": "<EVALUATOR_ID>" } }
  ```

  ```json Delete tasks variables theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
  { "input": { "onlineTaskIds": ["<TASK_ID>"] } }
  ```
</CodeGroup>

Reference: [`deleteEvaluator`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#deleteevaluator), [`deleteOnlineTask`](/docs/ax/graphql-reference/mutations/evaluators-and-tasks#deleteonlinetask).

## Gotchas and behavior notes

<AccordionGroup>
  <Accordion title="Several IDs are nullable in the schema despite being functionally required">
    `CreateEvalTaskMutationInput`/`PatchEvalTaskMutationInput.modelId` and `.datasetId` are plain `ID`, not `ID!`, and the schema does not enforce that you provide one or the other, so a task created with neither has no data source to scope against. `RunOnlineTaskMutationInput.onlineTaskId` is similarly typed `ID` even though there is no meaningful way to run a backfill without it.
  </Accordion>

  <Accordion title="EvalColumnName is a custom scalar, not a plain string or enum">
    Evaluator and column names (`TemplateEvaluationConfigInput.name`, `CodeEvaluationConfigInput.name`, and others) use the `EvalColumnName` scalar. The schema does not expose its validation rules; naming constraints are enforced server-side at request time.
  </Accordion>

  <Accordion title="Three different enums express the same maximize/minimize/none direction">
    Template evaluators use `TemplateEvaluationConfigDirection`, harness evaluators use `HarnessEvaluationConfigDirection`, and System One questions use `OptimizationDirection`. All three have the identical value set (`maximize`, `minimize`, `none`), but they are distinct enum types, so you cannot share a single constant across evaluator kinds in a strongly typed client.
  </Accordion>

  <Accordion title="System One's integration ID field is a String, not an ID">
    `SystemOneEvaluatorInput.llmIntegrationId` is typed `String!`, while the equivalent field on `TemplateEvaluationLlmConfigInput`, `HarnessEvaluatorInput`, and `OnlineTaskLLMConfigInput` is typed `ID!`. All of them expect the same relay global ID string; only the GraphQL type differs.
  </Accordion>

  <Accordion title="Several mutations in this domain have no description in the schema">
    `patchEvalTask`, `patchEvalTaskWithEvalHub`, `deleteOnlineTask`, and `patchPromptOptimizationTask` carry no top-level description, unlike their sibling `create` mutations. Field-level descriptions are present, but you have to infer each mutation's purpose from its name and input shape.
  </Accordion>
</AccordionGroup>

<CardGroup cols={3}>
  <Card title="Evaluator and task mutations reference" icon="book" href="/docs/ax/graphql-reference/mutations/evaluators-and-tasks">
    Full argument and field listing for all 14 evaluator and task mutations.
  </Card>

  <Card title="All GraphQL mutations" icon="list" href="/docs/ax/graphql-reference/mutations">
    Browse mutations for every other domain.
  </Card>

  <Card title="API explorer" icon="terminal" href="/docs/ax/graphql-reference/overview/api-explorer">
    Try queries and mutations interactively with autocomplete.
  </Card>
</CardGroup>
