NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2514 most downloaded on PyPI
The LLM Evaluation Framework
Last release 2 days ago
02 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Rarely documented
notes for 9 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
527 releases · first in 2023
One column per quarter.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
New to deepeval ? Get started here .
New to deepeval? Get started here.
This release also collects the changes from 4.1.2 through 4.1.10, which shipped to PyPI without a GitHub release.
4.1.104.1.84.1.84.1.64.1.34.1.34.1.104.1.64.1.54.1.34.1.24.1.104.1.104.1.104.1.104.1.104.1.104.1.84.1.84.1.54.1.34.1.34.1.34.1.34.1.34.1.2Contributors to these changes: @A-Vamshi, @Anai-Guo, @biztex, @chiruu12, @hassaanch23, @huisman, @jhochenbaum, @jsaimanoj, @Koustav-github, @kritinv, @lorenzozanee, @nightcityblade, @Patwaji, @penguine-ip, @Solaris-star, @tanayvaswani, @tonycueva, @tunglambk, @ykocaogullar
A huge thank you to everyone who contributed to this release ❤️
Nothing published for this version
New to deepeval ? Get started here .
New to deepeval? Get started here.
type in tool-calls metrics (#3062 by @tanayvaswani)New contributors: @bajiu-bajiu, @MarkHe1222
A huge thank you to everyone who contributed to this release ❤️
Nothing published for this version
New to deepeval? Get started here.
New to deepeval? Get started here.
New contributors: @eno42
A huge thank you to everyone who contributed to this release ❤️
Nothing published for this version
Nothing published for this version
Nothing published for this version
New to deepeval? Get started here.
New to deepeval? Get started here.
AgentLoopDetectionMetric for deterministic loop detection in agent traces. (#2645 by @Jeel3011)ToolPermissionMetric for metrics evaluation. (#2826 by @gh-raju)tools_called and relax type strictness. (#2864 by @A-Vamshi)capture_metric_type with the updated telemetry signature. (#2568 by @sachinML)tools_called_formatted into the extract_goal_and_outcome template. (#2808 by @ppcvote)UnboundLocalError when a document yields zero chunks. (#2849 by @hassaanch23)deepeval.evaluate.configs. (#2851 by @ErenAta16)New contributors: @maxtaran2010, @Jeel3011, @gh-raju, @sachinML, @ppcvote, @hassaanch23, @ErenAta16
A huge thank you to everyone who contributed to this release ❤️
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Add support for the claude-opus-4-8 model preset, including multimodal and structured output capabilities with updated pricing metadata. (#2698) (Vams
claude-opus-4-8 model preset, including multimodal and structured output capabilities with updated pricing metadata. (#2698) (Vamshi Adimalla)Nothing published for this version
ConversationSimulator now accepts simulation_graph, and controller is deprecated in favor of stopping_controller with a warning for legacy usage. (#26…
ConversationSimulator now accepts simulation_graph, and controller is deprecated in favor of stopping_controller with a warning for legacy usage. (#2678) (Jeffrey Ip)retrieval_context entries as RetrievedContextData with context and source, enabling contextual precision to group retrieved chunks by source. This also normalizes mixed inputs to plain context strings when preparing datasets and API payloads. (#2669) (Vamshi Adimalla)IMAGE and PDF placeholders. (#2672) (Vamshi Adimalla)A huge thank you to our contributors for this release: @penguine-ip, @A-Vamshi
DeepEval 4.0 introduces an agent-native evaluation workflow designed for coding agents, rapid debugging, and production AI systems.
DeepEval 4.0 introduces an agent-native evaluation workflow designed for coding agents, rapid debugging, and production AI systems.
If you're vibe coding agents, on something like claude code, this release is for you.
Coding agents can now run eval-driven iterations directly in context.
Quickstart: https://deepeval.com/docs/vibe-coder-quickstart
Inspect traces locally without leaving the terminal.
The idea is to develop rapidly by staying local on your machine. No more delegating to external UIs, unless collaboration is required.
<img width="1077" height="639" alt="Screenshot 2026-05-14 at 12 10 31 AM" src="https://github.com/user-attachments/assets/2309ad44-b423-46de-859e-5ce647fb14a3" />
One-line integrations for modern AI frameworks and providers.
Supports:
Here's a langchain example:
import pytest
from langchain.agents import create_agent
from deepeval import assert_test
from deepeval.integrations.langchain import CallbackHandler
from deepeval.dataset import EvaluationDataset, Golden
from deepeval.metrics import TaskCompletionMetric
def multiply(a: int, b: int) -> int:
return a * b
agent = create_agent(model="openai:gpt-4o-mini", tools=[multiply], system_prompt="Be concise.")
dataset = EvaluationDataset(goldens=[
Golden(input="What is 8 multiplied by 6?"),
Golden(input="What is 7 multiplied by 9?"),
])
@pytest.mark.parametrize("golden", dataset.goldens)
def test_langchain_agent(golden: Golden):
agent.invoke(
{"messages": [{"role": "user", "content": golden.input}]},
config={"callbacks": [CallbackHandler()]},
)
assert_test(golden=golden, metrics=[TaskCompletionMetric()])
That's it! Native CI/CD integration for LangChain agents, via Pytest, enabled by DeepEval.
Quickstart for integrations: https://deepeval.com/integrations
Nothing published for this version
If you're building agents, DeepEval can now analyze and give you metric scores based on the trace of your LLM app.
If you're building agents, DeepEval can now analyze and give you metric scores based on the trace of your LLM app.
Evaluate whether an agent actually completes the intended task, not just whether its final output “looks correct.”
Captures:
Docs: https://deepeval.com/docs/metrics-task-completion
Evaluates whether tools were invoked correctly, meaningfully, and in the right order.
Captures:
Docs: https://deepeval.com/docs/metrics-tool-correctness
Evaluates whether the agent’s arguments to tools are valid, structured, and aligned to the task.
Captures:
Docs: https://deepeval.com/docs/metrics-argument-correctness
Measures how efficiently an agent completes a task — rewarding fewer unnecessary steps and penalizing detours.
Captures:
Docs: https://deepeval.com/docs/metrics-step-efficiency
Evaluates how well the agent follows a predefined or self-generated plan.
Captures:
Docs: https://deepeval.com/docs/metrics-plan-adherence
Evaluates the quality of the plan itself when the agent generates one.
Captures:
Docs: https://deepeval.com/docs/metrics-plan-quality
Synthetic data generation now supports multi-turn goldens instead of just single-turn.
You can now generate:
Perfect for building large-scale synthetic datasets for support agents, sales agents, research assistants, workflow agents, and any multi-step conversational system.
from deepeval.synthesizer import Synthesizer
synthesizer = Synthesizer()
conversational_goldens = synthesizer.generate_conversational_goldens_from_docs(
document_paths=['example.txt', 'example.docx', 'example.pdf'],
)
Docs here (click on the "multi-turn" tab): https://deepeval.com/docs/synthesizer-generate-from-docs
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
If you're using any of the features below, you'll likely see a 50% reduction in code required, especially around ETL for formatting things in and out
If you're using any of the features below, you'll likely see a 50% reduction in code required, especially around ETL for formatting things in and out of DeepEval's ecosystem. This includes:
The first LLM-arena-as-a-Judge metric, now runs a blinded experiment and swaps positions randomly for a fair verdict on which LLM output is better.
Docs: https://deepeval.com/docs/metrics-arena-g-eval
Simply run your loop -> call your agent X number of times -> get your evaluation results. No more trying to fit non-test-case-friendly formats. Instead DeepEval will find your LLM traces automatically to run evals.
from somewhere import your_async_llm_app # Replace with your async LLM app
from deepeval.dataset import EvaluationDataset, Golden
dataset = EvaluationDataset(goldens=[Golden(input="...")])
for golden in dataset.evals_iterator():
# Create task to invoke your async LLM app
task = asyncio.create_task(your_async_llm_app(golden.input))
dataset.evaluate(task)
Docs: https://deepeval.com/docs/evaluation-component-level-llm-evals
Previously you have to define a list of user intentions, profile items, with a ton of more configs to juggle between. Now you can define a list of goldens with a standardized benchmark of scenarios to have turns generated for.
from deepeval.test_case import Turn
from deepeval.simulator import ConversationSimulator
# Create ConversationalGolden
conversation_golden = ConversationalGolden(
scenario="Andy Byron wants to purchase a VIP ticket to a cold play concert.",
expected_outcome="Successful purchase of a ticket.",
user_description="Andy Byron is the CEO of Astronomer.",
)
# Define chatbot callback
async def chatbot_callback(input):
return Turn(role="assistant", content=f"Chatbot response to: {input}")
# Run Simulation
simulator = ConversationSimulator(model_callback=chatbot_callback)
conversational_test_cases = simulator.simulate(goldens=[conversation_golden])
print(conversational_test_cases)
Docs: https://deepeval.com/docs/conversation-simulator
We also updated our docs with more improvements to come 👀
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →