NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2121 most downloaded on PyPI
The LLM Evaluation Framework
Last release today
29 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Rarely documented
notes for 9 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
526 releases · first in 2023
One column per quarter.
Nothing published for this version
Nothing published for this version
Automatically integrated with Confident AI for continous evaluation throughout the lifetime of your LLM (app):
Automatically integrated with Confident AI for continous evaluation throughout the lifetime of your LLM (app):
-log evaluation results and analyze metrics pass / fails -compare and pick the optimal hyperparameters (eg. prompt templates, chunk size, models used, etc.) based on evaluation results -debug evaluation results via LLM traces -manage evaluation test cases / datasets in one place -track events to identify live LLM responses in production -add production events to existing evaluation datasets to strength evals over time
Nothing published for this version
Nothing published for this version
Nothing published for this version
Automatically integrated with Confident AI for continous evaluation throughout the lifetime of your LLM (app):
Automatically integrated with Confident AI for continous evaluation throughout the lifetime of your LLM (app):
-log evaluation results and analyze metrics pass / fails -compare and pick the optimal hyperparameters (eg. prompt templates, chunk size, models used, etc.) based on evaluation results -debug evaluation results via LLM traces -manage evaluation test cases / datasets in one place -track events to identify live LLM responses in production -add production events to existing evaluation datasets to strength evals over time
Nothing published for this version
Nothing published for this version
Nothing published for this version
Mid-week bug fixes release with an extra feature:
Mid-week bug fixes release with an extra feature:
evaluate, evaluates a list of test cases (dataset) on metrics you define, all without having to go through the CLI. More info here: https://docs.confident-ai.com/docs/evaluation-datasets#evaluate-your-dataset-without-pytestIn this release, deepeval has added support for:
In this release, deepeval has added support for:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
ensure telemetry hits your server by @ColabDog in https://github.com/confident-ai/deepeval/pull/202
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.20.5...v0.20.6
firewall check for telemetry by @ColabDog in https://github.com/confident-ai/deepeval/pull/200
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.20.3...v0.20.5
Nothing published for this version
clean quickstart by @ColabDog in https://github.com/confident-ai/deepeval/pull/166
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.20.0...v0.20.3
Nothing published for this version
Nothing published for this version
Rename HOW_TO_CONTRIBUTE.md to CONTRIBUTING.md by @penguine-ip in https://github.com/confident-ai/deepeval/pull/164
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.19.0...v0.20.0
Nothing published for this version
Nothing published for this version
add guardrails integration by @ColabDog in https://github.com/confident-ai/deepeval/pull/158
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.18.0...v0.19.0
Add new customer support example by @ColabDog in https://github.com/confident-ai/deepeval/pull/154
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.17.9...v0.18.0
fix by @ColabDog in https://github.com/confident-ai/deepeval/pull/147
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.17.6...v0.17.8
Nothing published for this version
adding length metric including test and documentation by @j-space-b in https://github.com/confident-ai/deepeval/pull/139
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.17.5...v0.17.6
Hotfix/fix conceptual similarity threshold by @ColabDog in https://github.com/confident-ai/deepeval/pull/133
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.17.4...v0.17.5
Feature/cache pip dependencies by @ColabDog in https://github.com/confident-ai/deepeval/pull/128
Full Changelog: https://github.com/confident-ai/deepeval/compare/v0.17.3...v0.17.4
Feature/add synthetic query generation by @ColabDog in https://github.com/confident-ai/deepeval/pull/1
Full Changelog: https://github.com/confident-ai/deepeval/commits/v0.17.3
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →