Skip to main content
Edit a dataset’s examples directly in the table and save them as one new version. Trace tree rows now show latency, tokens, and cost, and Phoenix serves your own Agent Skills over MCP and in PXI.

Breaking Changes in the GraphQL API

September 16, 2026 Breaking change in arize-phoenix 20.13.0
  • rootSpansOnly and orphanSpanAsRootSpan are removed from Project.spans and Trace.spans. Use filterCondition: "parent_id is None" instead, or "parent_span is None" to count orphans as roots.
  • Project.traceAnnotationsNames is now Project.traceAnnotationNames.
  • applyDatasetExampleChanges is now patchDatasetExamples. It takes an ordered, JSON Patch-style list of operations that commits as one dataset version, and requires datasetId.

Edit Dataset Examples in Place

September 16, 2026 Available in arize-phoenix 20.13.0+ Click Edit above the examples table to fix input, output, and metadata JSON without opening examples one at a time. The whole batch saves as one new dataset version.
  • Edit JSON cells in place, add rows with an optional custom ID, and remove or restore rows before you save. Phoenix flags duplicate custom IDs before the save goes through.
  • Save with ⌘S. The toolbar counts pending changes, the save dialog takes an optional version description, and Phoenix warns before you leave the page with unsaved edits.
  • Search and split filters stay live while you edit, and the toolbar reports any changes the current search hides.

Updating Datasets

Ways to change an existing dataset

Splits

Group examples into named subsets

Latency, Tokens, and Cost on Every Trace Tree Row

September 22 to 23, 2026 Available in arize-phoenix 20.16.0+ Each trace tree row now ends with latency, tokens, and cost, so you can spot the slow or expensive span without opening it.
  • Hover a row to preview its breakdown. One preview follows the pointer down the tree and loads details only once the pointer settles, so large traces stay responsive.
  • Tokens and cost share one breakdown. Each measure splits into token types, so you can compare a type’s share of tokens with its share of cost.

PXI: Custom Skills, Schema Search, and Slash-Command Keys

September 21 to 29, 2026 Available in arize-phoenix 20.15.0+ (server) and @arizeai/phoenix-cli 1.18.6+ Point Phoenix at your own Agent Skills and it serves them next to the built-in ones, both in the MCP server’s instructions and in PXI’s skill picker.
  • Each path is a skill or a directory of skills. A skill’s directory name must match the name in its SKILL.md frontmatter, and the startup banner lists what mounted. Use absolute paths in deployments.
  • PXI searches the GraphQL schema instead of reading the full SDL: phoenix-gql schema --search <terms> returns ranked field signatures, and phoenix-gql schema --names <Type.field> prints one definition and the paths that reach it.
  • Slash-command hints respond to the keyboard in the pxi terminal client. Arrow keys move the highlight, Tab completes, and Enter runs, so /he plus Enter runs /help.

Remote MCP Server

Mount your own skills and configure the endpoint

PXI

Investigate Phoenix data with PXI

Score Classifications with Decision-Only Models

September 21, 2026 Available in @arizeai/phoenix-evals 2.6.0+ (TypeScript) Pass an AI SDK evaluation model, such as TypeSafe’s Jev, to any classification evaluator. Evaluation models answer one typed choice question instead of generating text, so a single-label eval runs faster and costs less than on a chat model.
  • Same evaluators, same labels. Phoenix detects the evaluation model and sends the call through experimental_evaluate instead of generateObject.
  • Results carry a label and score but no explanation. metadata reports the full label distribution as probabilities, along with the modelId.

TypeSafe AI

Trace and evaluate with TypeSafe models

Pre-Built Metrics

Every evaluator that ships with Phoenix

GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5

September 23, 2026 Available in arize-phoenix 20.16.0+ All three models are available in the playground and priced in the built-in cost manifest, including the 272k long-context tier for the GPT-6 pair.
  • GPT-6 Sol and GPT-6 Luna are reasoning models, so Phoenix sends them to the Responses API on OpenAI and Azure OpenAI, where tool calls work.
  • Claude Opus 5.5 always uses adaptive thinking, so the playground omits temperature, top-p, top-k, and extended-thinking settings. Bedrock lists it as anthropic.claude-opus-5-5.
  • Opus 5.5 spans bill at Opus 5.5 rates. They previously matched the claude-opus-5 pattern and billed at Opus 5’s higher rates.

TypeScript Client: Secrets, Splits, Prompt Tags, and Filters

September 21 to 29, 2026 Available in @arizeai/phoenix-client 7.12.0+ New helpers manage secrets, dataset splits, and prompt version tags without a hand-written REST call, and trace and session listing now accept filter expressions.
  • upsertOrDeleteSecrets (7.13.0) applies a batch atomically: a value creates or updates a key, and null deletes it. The response lists key names only, never values.
  • createDatasetSplit, updateDatasetSplit, and deleteDatasetSplit (7.14.0) manage splits on a dataset selected by name or GlobalID. Requires Phoenix server 19.20.0+.
  • upsertPromptVersionTag and deletePromptVersionTag (7.15.0) create, move, and remove prompt version tags.
  • filter on getTraces and listSessions (7.12.0) takes the same expressions as the UI and deprecates the error, minLatencyMs, and maxLatencyMs parameters. Requires Phoenix server 20.12.0+.

Filter Expressions

Filter vocabulary for spans, traces, and sessions

Tag a Prompt

Promote a prompt version with a tag

REST API Updates

September 23, 2026 Available in arize-phoenix 20.16.0+
  • GET /v1/projects/{project_identifier}/spans accepts sort=start_time, so a page limit returns the latest spans rather than the most recently ingested ones. The default sort=id keeps insertion order.

Additional Improvements

September 15 to 29, 2026 Available in arize-phoenix 20.13.0+, arize-phoenix-evals 3.9.0+, and arize-phoenix-otel 0.17.2+
  • Experiment names and descriptions are editable from the experiments table and when you launch an experiment from the playground.
  • register() no longer crashes with opentelemetry-exporter-otlp-proto-http 1.45.
  • Analytics SQL on SQLite supports the bundled time functions (time_trunc, time_parse, time_fmt_datetime, and time_sub), and latency_ms is nanosecond-exact.
  • PXI streams faster, keeps partial output when you interrupt a turn, and closes the turn’s trace when a server tool raises.
  • PromptTemplate message lists accept role aliases such as developer, human, ai, and model.
  • The rate limiter rejects an invalid initial request rate up front.
  • SQLite autoincrement counters survive migrations that rebuild a table, so upgrades never reuse a primary key.
  • Metric chart time ranges stay stable, and the table’s load-more control stays in the footer in Safari.
  • The cost manifest carries refreshed token prices for built-in models.