NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2514 most downloaded on PyPI
The LLM Evaluation Framework
Last release 2 days ago
02 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Rarely documented
notes for 9 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
527 releases · first in 2023
One column per quarter.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
https://www.deepeval.com/docs/metrics-llm-evals#rubric
https://www.deepeval.com/docs/metrics-llm-evals#rubric
Nothing published for this version
Nothing published for this version
Nothing published for this version
In this release we've cleaned up some dependencies to separate out dev packages, as well as more tracing verbose logs for debugging.
In this release we've cleaned up some dependencies to separate out dev packages, as well as more tracing verbose logs for debugging.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
> ⚠️ This release introduces breaking changes in preparation for DeepEval v3.0.
⚠️ This release introduces breaking changes in preparation for DeepEval v3.0. Please review carefully and adjust your code as needed.
evaluate() function now has "configs"evaluate() function had 13+ arguments to control display, async behaviors, caching, etc. and it was growing out of control. We've now abstracted it into "configs" instead:from deepeval.evaluate.configs import AsyncConfig
from deepeval import evaluate
evaluate(..., async_config=AsyncConfig(max_concurrent=20))
Full docs here: https://www.deepeval.com/docs/evaluation-running-llm-evals#configs-for-evaluate
This shouldn't be a surprised but, DeepTeam now takes care of everything red teaming related, for the foreseeable future. Docs here: https://trydeepteam.com
Nested components are a mess to evaluate. In this version in preparation for v3.0, we introduced dynamic evals, where you can apply a different set of metrics for different components in your LLM application:
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.tracing import observe, update_current_span_test_case
@observe(metrics=[AnswerRelevancyMetric()])
def complete(query: str):
response = openai.ChatCompletion.create(model="gpt-4o", messages=[{"role": "user", "content": query}]).choices[0].message["content"]
update_current_span_test_case(
test_case=LLMTestCase(input=query, output=response)
)
return response
Full docs here: https://www.deepeval.com/docs/evaluation-running-llm-evals#setup-tracing-highly-recommended
Nothing published for this version
Cleaned up dependencies for upcoming 3.0 release:
Cleaned up dependencies for upcoming 3.0 release:
Removed the automatic updates, it is now opt-in: https://www.deepeval.com/docs/miscellaneous
Removed instructor, double checked and it wasn't used anywhere
Removed LlamaIndex and moved it to optional, only needed for one module
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
The latest conversation simulator simulates fake user interactions to generate conversations on your behalf. These conversations can be used for evalu
The latest conversation simulator simulates fake user interactions to generate conversations on your behalf. These conversations can be used for evaluation right afterwards, and is similar to the goldens synthesizer. Docs here: https://docs.confident-ai.com/docs/evaluation-conversation-simulator
Nothing published for this version
Nothing published for this version
Migrated default provider models to support Synthesizer
What's New 🔥
deepeval < 2.5.6 might need to update importsNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Custom prompt template overriding for all RAG metrics. This was introduced for folks using weaker models for evaluation, or just models in general tha
deepeval's metrics are built around. You can still use your favorite metrics and algorithms, but now with a custom template if required. Example here: https://docs.confident-ai.com/docs/metrics-answer-relevancy#customize-your-templatesave_as() for datasets to save test cases as well: https://docs.confident-ai.com/docs/evaluation-datasets#save-your-datasetSynthesizerDAGMetric: https://docs.confident-ai.com/docs/metrics-dagNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
🥳 Latest feature to allow users to inject the Faithfulness metric with their custom template. Most suited for custom LLMs where text data is highly fo
🥳 Latest feature to allow users to inject the Faithfulness metric with their custom template. Most suited for custom LLMs where text data is highly formatted by data engineers and stored in databases according to different categories.
Your coding agent can read these notes before it upgrades. Set up the MCP server →