NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3171 most downloaded on PyPI
AI Observability and Evaluation
Last release today
16 Sep 2026
Ships on a steady schedule
a new release about every 8 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
19 versions withdrawn
withdrawn after publishing
4 years old
694 releases · first in 2023
One column per quarter.
agents: add PHOENIX_AGENTS_DISABLE_BASH to disable server-side bash
resolve route path through FastAPI 0.137 _IncludedRouter in prometheus middleware
agents: experiment editing & eval skills
agent: add a subagents toggle to assistant settings
agents: Add local slash commands to chat menu
add playground repetitions tool
playground: add PXI load_dataset tool
server: add system settings with admin-managed assistant enablement and trace recording policy
add PXI playground model switching tool
add playground save prompt tool
agent: collapse agent context pills into a deck
agents: adds skill for creating a new eval dataset, a new dat...
Phoenix now lets you compose evaluation strategies in code.
Phoenix now lets you compose evaluation strategies in code.
Most eval tooling hands you a fixed menu of judge templates. Real evaluation is rarely that tidy.
Code Evaluators enable you to build evaluation criteria the way you want. You write a Python or TypeScript evaluate() function in the Phoenix UI — no SDK, no local runtime, no deploy step — and Phoenix runs it server-side, recording labels and scores as annotations on every experiment run.
Because it's just code, you control the whole strategy:
• Composite scoring: blend sub-scores (LLM judgment + deterministic rules) into one weighted metric • Embedding-based evaluation: cosine similarity over embeddings instead of brittle string matching • LLM juries: poll multiple models and combine verdicts into a weighted consensus
Sandboxed Code evaluators unlock the idea of agents as a judge as well. We're excited where this is heading.
agents: stamp tool execution environment on tool-call provider metadata
agents: instrument PXI agent system prompt on LLM spans
agents: Generative UI rendering within agent chats
agents: persist agent turn as a root span
add trace feedback toolbar to session turns
agents: populate project_sessions for /chat and /summary traces
add trace user feedback annotations
agents: agent set_time_range tool with hardened context injection
agent: remove user instructions from PXI
add TanStack AI tracing integration
api: query annotations by identifier on GET endpoints
extend db prompt types for invocation paramaters
agents: support custom providers and secret store values
agents: advertise phoenix page context to chat
cli: add trace note support to px
Add a dedicated span notes column and clean up annotation selection
agent: session summary eval + sidebar-title prompt rewrite
agent: Display token count for chat sessions
db: wrap Azure AsyncEngine with ObjectProxy instead of patching dispose
Add PXI consent and trace-sharing controls
db: add Azure managed identity authentication for PostgreSQL
ui: refresh projects page on mount and on 60s interval
agent: fix docs URL formatting in PXI system prompt + add eval harness
Add chat metadata to assistant messages
agent: Capability menu + Refactor PXI frontend tool wiring around a capability registry
Harden authFetch and add it to dataset upload path
remove WebSocket support for GraphQL subscriptions, use HTTP multipart
add /redirects/projects/:project_name for name-based project URLs
use dataloader for session annotations to prevent connection exhaustion
add valid OpenAPI examples for span creation endpoint
evals: deprecate evals 1.0 and remove legacy experiments module
See MIGRATION.md
use DISTINCT instead of GROUP BY in ProjectHasTracesDataLoader
rest-api: DELETE /prompt_versions/{id}/tags/{tag_name}
ui: fix model menu search losing focus on first keystroke
Your coding agent can read these notes before it upgrades. Set up the MCP server →