NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2121 most downloaded on PyPI
The LLM Evaluation Framework
Last release 3 days ago
14 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Rarely documented
notes for 9 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
522 releases · first in 2023
One column per quarter.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
DeepEval's 3.2.6 release focuses on single-vs multi-turn use cases in datasets!
DeepEval's 3.2.6 release focuses on single-vs multi-turn use cases in datasets!
input → output pairs for one-off prompt testing.DeepEval now automatically detects whether a dataset is single-turn or multi-turn based on structure and routes to the appropriate evaluation logic.
Introduced a new concept: conversational goldens, which contains scenario, (and optionally expected_outcome) but not things like input and expected output as with single-turn use cases..
This release is setting the stage for future multi-turn use cases.
Docs here: https://deepeval.com/docs/evaluation-datasets
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
In DeepEval's latest release, we are introducing ArenaGEval, the first ever metric to compare test cases to choose the best performing one based on yo
In DeepEval's latest release, we are introducing ArenaGEval, the first ever metric to compare test cases to choose the best performing one based on your custom criteria.
It looks something like this:
from deepeval import evaluate
from deepeval.test_case import ArenaTestCase, LLMTestCaseParams
from deepeval.metrics import ArenaGEval
a_test_case = ArenaTestCase(
contestants={
"GPT-4": LLMTestCase(
input="What is the capital of France?",
actual_output="Paris",
),
"Claude-4": LLMTestCase(
input="What is the capital of France?",
actual_output="Paris is the capital of France.",
),
},
)
arena_geval = ArenaGEval(
name="Friendly",
criteria="Choose the winter of the more friendly contestant based on the input and actual output",
evaluation_params=[
LLMTestCaseParams.INPUT,
LLMTestCaseParams.ACTUAL_OUTPUT,
],
)
metric.measure(a_test_case)
print(metric.winner, metric.reason)
Docs here: https://deepeval.com/docs/metrics-arena-g-eval
Nothing published for this version
Nothing published for this version
Nothing published for this version
Previously we had great support for single-turn, text evaluation in the form of LLMTestCases, but now we're adding MLLMTestCase, which accepts images:
Previously we had great support for single-turn, text evaluation in the form of LLMTestCases, but now we're adding MLLMTestCase, which accepts images:
from deepeval.metrics import MultimodalGEval
from deepeval.test_case import MLLMTestCaseParams, MLLMTestCase, MLLMImage
from deepeval import evaluate
m_test_case = MLLMTestCase(
input=["Show me how to fold an airplane"],
actual_output=[
"1. Take the sheet of paper and fold it lengthwise",
MLLMImage(url="./paper_plane_1", local=True),
"2. Unfold the paper. Fold the top left and right corners towards the center.",
MLLMImage(url="./paper_plane_2", local=True)
]
)
text_image_coherence = MultimodalGEval(
name="Text-Image Coherence",
criteria="Determine whether the images and text is coherence in the actual output.",
evaluation_params=[MLLMTestCaseParams.ACTUAL_OUTPUT],
)
evaluate(test_cases=[m_test_case], metrics=[text_image_coherence])
Docs here: https://deepeval.com/docs/multimodal-metrics-g-eval
PS. This also includes platform support
<img width="1728" alt="Screenshot 2025-06-19 at 3 46 12 PM" src="https://github.com/user-attachments/assets/b05166be-d407-4a17-8af0-2e86dbdf72b2" />
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Previously we assumed a conversation as as a list of LLMTestCases, which might necessarily be the case. Now a conversational test case is made up of a
Previously we assumed a conversation as as a list of LLMTestCases, which might necessarily be the case. Now a conversational test case is made up of a list of Turns instead, which follows OpenAI's standard messages format:
from deepeval.test_case import Turn
turns = [Turn(role="user", content="...")]
Docs here: https://deepeval.com/docs/evaluation-test-cases#conversational-test-case
Nothing published for this version
Added new loading bars for component-level evals, and deepeval view to see results on Confident AI.
Added new loading bars for component-level evals, and deepeval view to see results on Confident AI.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →