NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2121 most downloaded on PyPI
The LLM Evaluation Framework
Last release 4 days ago
24 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Rarely documented
notes for 9 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
525 releases · first in 2023
One column per quarter.
Nothing published for this version
Nothing published for this version
Nothing published for this version
asynchronous support throughout deepeval, and no longer using threads. Users can also call individual metrics asynchronously: https://docs.confident-a
In deepeval v0.20.85:
evaluate() function for more customizability: https://docs.confident-ai.com/docs/evaluation-introduction#evaluating-without-pytestNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
In DeepEval's latest release, there is now:
In DeepEval's latest release, there is now:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
For the newest release, deepeval now is now stable for production use:
For the newest release, deepeval now is now stable for production use:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
For the latest release, DeepEval:
For the latest release, DeepEval:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
LLM-Evals (LLM evaluated metrics) now support all of langchain's chat models.
LLMTestCase now has execution_time and cost, useful for those looking to evaluate on these parametersminimum_score is now threshold instead, meaning you can now create custom metrics that either have a "minimum" or "maximum" thresholdLLMEvalMetric is now GEvalNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Faithfulness, Answer Relevancy, Contextual Relevancy, Contextual Precision, and Contextual Recall, all offer a reasoning for its given score.
In this release:
transformers, sentence_transformers, and pandas to reduce package sizeNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Lots of new features this release:
Lots of new features this release:
JudgementalGPT now allows for different languages - useful for our APAC and European friendsRAGAS metrics now supports all OpenAI models - useful for those running into context length issuesLLMEvalMetric now returns a reasoning for its scoredeepeval test run now has hooks that call on test run completionevaluate now displays retrieval_context for RAG evaluationRAGAS metric now displays metric breakdown for all its distinct metricsNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →