NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4785 most downloaded on PyPI
AI Observability and Evaluation
Last release 2 days ago
01 Oct 2026
Ships on a steady schedule
a new release about every 8 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
19 versions withdrawn
withdrawn after publishing
4 years old
700 releases · first in 2023
One column per quarter.
Deprecate datasets module, rename to inferences
⚠ BREAKING CHANGES
add bedrock claude tracing tutorial
ensure recent version of opentelemetry-proto is used
Add log_traces method that sends TraceDataset traces to Phoenix
Add SQL and Code Functionality Eval Templates
ui: show formatted JSON for attributes
Nothing published for this version
ui: broken context for markdown
UI: color rotation for markdown
gql: add trace node and trace evaluations
ignore docs/ directory when formatting
long project names do not overflow and squash project icon
Add response_format argument to MistralAIModel
trace: query dsl for numpy arrays
trace: eliminate truth ambiguity with non-empty numpy arrays
### Features * delete project ui
handle numpy types in json.dumps for gql
launch_app() with experimental span storage using environment variables for storage path and storage type enums
increase attributes limit on spans
### Bug Fixes * sanitize base path
experimental span storage with append-only text files
ui: scroll column selector when long
add arize-phoenix support for python 3.12
Enable dynamic project switching
display newlines in explanations
graphql: embed project inside graphql span as private attribute
projects: add support for the PHOENIX_PROJECT_NAME param
evals: deprecate document relevance evaluators
evals: deprecate document relevance evaluators
The Phoenix evals module is graduating out of experimental! You can now install Phoenix evals as a standalone package with pip install arize-phoenix-evals or you can include the new version of phoenix.evals along with the Phoenix install with pip install -U arize-phoenix[evals]. Swapping to the new evals module includes a few small breaking changes which might require some migration work. Details can be found in MIGRATION.md.
phoenix.experimental.evals is being deprecated and will remain in Phoenix for about a month before being removed.
phoenix.evals into phoenix (#2420) (dd3e7b4)evals: deprecate document relevance evaluators
evals: add session-level pii_detection evaluator
evals: add retrieval relevance evaluator
phoenix.evals (#2421) (fbd4961)BedrockModel (#2425) (81a720c)remove symbolic links for docker build
server: GET /v1/model_providers no longer returns custom providers or a next_cursor , and built-in entries expose provider rather than kind + provider
next_cursor, and built-in entries expose provider rather than kind + provider_key. Custom providers move to GET /v1/custom_model_providers. The endpoint is unreleased, so no published client is affected.pxi: enable phoenix-gql mutations by default with approval in manual mode
llama_index_search_and_retrieval_notebook
server: add POST /traces/transfer
fix: cast message to string in vertexai model
release arize-phoenix-client 3.1.0
trace: perform library version compatibility on llama_index
run_evals correctly falls back to default responses on error
handle ndarray during ingestion
client: replace google-generativeai formatter with google-genai
ui: add last_hour, fix end of hour rounding
evals: properly use kw args for models in notebooks
embeddings: add search by text and ID on selection
disregard active session if endpoint is provided to px.Client
absolute path for eval exporter
localhost address for px.Client
absolute path for urljoin in px.Client
phoenix client get_evaluations() and get_trace_dataset()
Remove model-level tenacity retries
persistence: add a PHOENIX_WORKING_DIR env var for setting up a…
b7b7dfb : Classification evaluators now accept AI SDK evaluation models such as TypeSafe's Jev. When an evaluation model is passed as model , the clas
model, the classification is routed through experimental_evaluate as a single choice question instead of generateObject. Results carry a label and score but no explanation, since evaluation models do not generate text.d67ea3f : Deprecate createDocumentRelevanceEvaluator in favor of createRetrievalRelevanceEvaluator . Rename documentText to context and the unrelated…
createCompletenessEvaluator to judge whether every active user request in a conversation was actually completed.createDocumentRelevanceEvaluator in favor ofcreateRetrievalRelevanceEvaluator. Rename documentText to context and theunrelated label to irrelevant. The deprecated factory and its types will beYour coding agent can read these notes before it upgrades. Set up the MCP server →