NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4896 most downloaded on PyPI
LLM framework to build customizable, production-ready LLM applications. Connect components (models, vector DBs, file converters) to pipelines or agents that can interact with your data.
Last release today
18 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
210 releases · first in 2023
Nothing published for this version
Two new retriever components make it easier to build hybrid search pipelines. MultiRetriever runs multiple text retrievers in parallel and merges thei
MultiRetriever and TextEmbeddingRetrieverTwo new retriever components make it easier to build hybrid search pipelines. MultiRetriever runs multiple text retrievers in parallel and merges their results into a single deduplicated list, ranked by reciprocal rank fusion by default. You can selectively enable or disable individual retrievers at runtime using the active_retrievers parameter. This is useful when you want to skip the embedding retriever for short or keyword-only queries, for example.
TextEmbeddingRetriever wraps an embedding-based retriever together with a text embedder into a single component, making it compatible with MultiRetriever by implementing the TextRetriever protocol. Here's how to combine BM25 and embedding retrieval in a single component:
from haystack.components.retrievers import MultiRetriever, TextEmbeddingRetriever
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever, InMemoryEmbeddingRetriever
from haystack.components.embedders import SentenceTransformersTextEmbedder
retriever = MultiRetriever(
retrievers={
"bm25": InMemoryBM25Retriever(document_store=doc_store),
"embedding": TextEmbeddingRetriever(
retriever=InMemoryEmbeddingRetriever(document_store=doc_store),
text_embedder=SentenceTransformersTextEmbedder(model="sentence-transformers/all-MiniLM-L6-v2"),
),
},
top_k=3,
)
# Run all retrievers
result = retriever.run(query="green energy sources")
# Run only the BM25 retriever
result = retriever.run(query="green energy sources", active_retrievers=["bm25"])
LLM.run and LLM.run_async no longer accept messages and streaming_callback as positional arguments — they must now be passed as keyword arguments. Update any direct calls accordingly:
# Before
llm.run([message], my_callback)
# After
llm.run(messages=[message], streaming_callback=my_callback)
run_async to CacheChecker, enabling it to be used in AsyncPipeline without blocking the event loop.Pipeline.connect(). When multiple senders are connected to the same list-typed receiver socket, ordering depends on the pipeline class. With Pipeline, items are ordered alphabetically by sender component name (because Pipeline.run() schedules components in alphabetical order for deterministic execution), not by the order of connect() calls. With AsyncPipeline, no ordering is guaranteed, since components in different branches may run in parallel. The docstrings now point users to a dedicated joiner component when they need explicit ordering.join_mode parameter to the experimental MultiRetriever component, supporting "reciprocal_rank_fusion" (default) and "concatenate". Reciprocal Rank Fusion merges the ranked result lists from all retrievers into a single deduplicated list ordered by RRF score. The underlying RRF logic is extracted into a shared utility _reciprocal_rank_fusion in haystack.utils.misc, which is now also used by DocumentJoiner.LLM now supports two usage modes:
user_prompt with Jinja2 variables (e.g. {{ query }}).
Those variables become pipeline inputs and messages is optional. The rendered user_prompt
is always appended after any messages provided at runtime.user_prompt or provide one with no template variables. messages
becomes a required input, allowing a fully-constructed list of ChatMessages to be passed from upstream.NamedEntityExtractor where the spaCy/Thinc device state was not correctly restored after execution, potentially affecting the device configuration of other spaCy components in the same process.chat_generator and tool_invoker) and to pipeline-level inputs, original_input_data, and pipeline_outputs captured by _create_pipeline_snapshot. When every field fails to serialize, the snapshot still stores a structurally valid empty payload ({"serialization_schema": {"type": "object", "properties": {}}, "serialized_data": {}}) so that resuming the snapshot does not raise DeserializationError — for example when resuming from a ToolBreakpoint where the sub-component's inputs are not strictly required.tools_strict=True in OpenAIChatGenerator to recursively apply additionalProperties: false and required to all nested objects in tool parameter schemas. Previously only the top-level object was transformed, causing OpenAI's strict mode to reject tools with nested parameters.@Aftabbs, @albertodiazdurana, @anakin87, @ArkaD171717, @bilgeyucel, @bogdankostic, @davidsbatista, @FuturMix, @julian-risch, @kacperlukawski, @ritikraj2425, @saivedant169, @shaun0927, @sjrl, @SyedShahmeerAli12
One column per month.
Nothing published for this version
Nothing published for this version
As part of the migration from requests to httpx, request_with_retry and async_request_with_retry (in haystack.utils.requests_utils) no longer raise re
As part of the migration from requests to httpx, request_with_retry and async_request_with_retry (in haystack.utils.requests_utils) no longer raise requests.exceptions.RequestException on failure; they now raise httpx.HTTPError instead. This also affects HuggingFaceTEIRanker, which relies on these utilities. Users catching requests.exceptions.RequestException should update their code to catch httpx.HTTPError.
The LLM component now requires user_prompt to be provided at initialization and it must contain at least one Jinja2 template variable (e.g. {{ variable_name }}). This ensures the component always exposes at least one required input socket, which is necessary for correct pipeline scheduling.
required_variables now defaults to "*" (all variables in user_prompt are required), and passing an empty list raises a ValueError.
If you are affected: update any code that instantiates LLM without a user_prompt, or with a user_prompt that has no template variables, to include at least one variable.
Before:
llm = LLM(chat_generator=OpenAIChatGenerator(), system_prompt="You are helpful.")
After:
llm = LLM(
chat_generator=OpenAIChatGenerator(),
system_prompt="You are helpful.",
user_prompt='{% message role="user" %}{{ query }}{% endmessage %}',
)
Agent.run() and Agent.run_async() now require messages as an explicit argument (no longer optional). If you were relying on the default None value in Haystack version 2.26 or 2.27, pass an empty list instead:
agent.run(messages=[], ...)
LLM.run() and LLM.run_async() are unaffected — they still accept None and default to an empty list internally.
Tools and components can now declare a State (or State | None) parameter in their signature to receive the live agent State object at invocation time — no extra wiring needed.
For function-based tools created with @tool or create_tool_from_function, add a state parameter annotated as State:
from haystack.components.agents import State
from haystack.tools import tool
@tool
def my_tool(query: str, state: State) -> str:
"""Search using context from agent state."""
history = state.get("history")
...
For component-based tools created with ComponentTool, declare a State input socket on the component's run method:
from haystack import component
from haystack.components.agents import State
from haystack.tools import ComponentTool
@component
class MyComponent:
@component.output_types(result=str)
def run(self, query: str, state: State) -> dict:
history = state.get("history")
...
tool = ComponentTool(component=MyComponent())
In both cases ToolInvoker automatically injects the runtime State object before calling the tool, and State/Optional[State] parameters are excluded from the LLM-facing schema so the model is not asked to supply them.
This is an alternative to the existing inputs_from_state and outputs_to_state options on Tool and ComponentTool, which map individual state keys to specific tool parameters and outputs declaratively. Injecting the full State object is more flexible and useful when a tool needs to read from or write to multiple keys, but it couples the tool implementation directly to State.
DocumentCleaner with its default settings can flatten Markdown output, and update the example pipelines for PaddleOCRVLDocumentConverter, MistralOCRDocumentConverter, AzureDocumentIntelligenceConverter, and MarkItDownConverter to avoid routing Markdown content through the default cleaner configuration._create_agent_snapshot robust towards serialization errors. If serializing agent component inputs fails, a warning is logged and an empty dictionary is used as a fallback, preventing the serialization error from masking the real pipeline runtime error.httpx for both synchronous and asynchronous requests, replacing requests. Error reporting for failed requests has also been improved: exceptions now include additional details alongside the reason field.run_async method to LLMMetadataExtractor. ChatGenerator requests now run concurrently using the existing max_workers init parameter.MarkdownHeaderSplitter now accepts a header_split_levels parameter (list of integers 1–6, default all levels) to control which header depths create split boundaries. For example, header_split_levels=[1, 2] splits only on # and ## headers, merging content under deeper headers into the preceding chunk.MarkdownHeaderSplitter now ignores # lines that appear inside fenced code blocks (triple-backtick or triple-tilde), preventing Python comments and other hash-prefixed lines in code from being misidentified as Markdown headers.PaddleOCRVLDocumentConverter documentation with more detailed guidance on advanced parameters, common usage scenarios, and a more realistic configuration example for layout-heavy documents.Fix ToolInvoker._merge_tool_outputs silently appending None to list-typed state when a tool's outputs_to_state source key is absent from the tool result. This is a common scenario with PipelineTool wrapping a pipeline that has conditional branches where not all outputs are always produced even if defined in outputs_to_state. The mapping is now skipped entirely when the source key is not present in the result dict.
When using the MarkdownHeaderSplitter, in the split chunks, the child header previously lost its direct parent header in the metadata. Previously if one executed the code below:
from haystack.components.preprocessors import MarkdownHeaderSplitter
from haystack import Document
text = """
# header 1
intro text
## header 1.1
text 1
## header 1.2
text 2
### header 1.2.1
text 3
### header 1.2.2
text 4
"""
document = Document(content=text)
splitter = MarkdownHeaderSplitter(
keep_headers=True,
secondary_split="word"
)
result = splitter.run(documents=[document])["documents"]
for doc in result:
print(f"Header: {doc.meta['header']}, parent headers: {doc.meta['parent_headers']}")
We would have expected this output:
Header: header 1, parent headers: []
Header: header 1.1, parent headers: ['header 1']
Header: header 1.2, parent headers: ['header 1']
Header: header 1.2.1, parent headers: ['header 1', 'header 1.2']
Header: header 1.2.2, parent headers: ['header 1', 'header 1.2']
But instead we actually got:
Header: header 1, parent headers: []
Header: header 1.1, parent headers: []
Header: header 1.2, parent headers: ['header 1']
Header: header 1.2.1, parent headers: ['header 1']
Header: header 1.2.2, parent headers: ['header 1', 'header 1.2']
The error happened when a parent header had its own content chunk before the first child header.
This has been fixed so even when a parent header has its own content chunk before the first child header all content is preserved.
Reverts the change that made Agent messages optional as it caused issues with pipeline execution. As a consequence, the LLM component now defaults to an empty messages list unless provided at runtime.
@Aftabbs, @Amanbig, @anakin87, @bilgeyucel, @bogdankostic, @davidsbatista, @dina-deifallah, @jimmyzhuu, @julian-risch, @kacperlukawski, @maxdswain, @MechaCritter, @ritikraj2425, @sarahkiener, @sjrl, @soheinze, @srini047, @tholor
Nothing published for this version
Nothing published for this version
When a component expects a list as input, pipelines now automatically join multiple inputs into that list (no extra components needed), even if they c
When a component expects a list as input, pipelines now automatically join multiple inputs into that list (no extra components needed), even if they come in different but compatible types. This enables patterns like combining a plain query string with a list of ChatMessage objects into a single list[ChatMessage] input.
Supported conversations:
| Source Types | Target Type | Behavior |
|---|---|---|
| T + T | list[T] | Combines multiple inputs into a list of the same type. |
| T + list[T] | list[T] | Merges single items and lists into a single list. |
| str + ChatMessage | list[str] | Converts all inputs to str and combines into a list. |
| str + ChatMessage | list[ChatMessage] | Converts all inputs to ChatMessage and combines into a list. |
Learn more about how to simplify list joins in pipelines in 📖 Smart Pipeline Connections: Implicit List Joining
The metadata inspection and filtering utilities (count_documents_by_filter, count_unique_metadata_by_filter, get_metadata_field_min_max, etc.) are now available in the InMemoryDocumentStore, aligning it with other document stores.
You can prototype locally in memory and easily debug, filter, and inspect the data in the document store during development, then reuse the same logic in production. See all available methods in InMemoryDocumentStore API reference.
Added new operations to the InMemoryDocumentStore: count_documents_by_filter, count_unique_metadata_by_filter, get_metadata_fields_info, get_metadata_field_min_max, get_metadata_field_unique_values
AzureOpenAIChatGenerator now exposes a SUPPORTED_MODELS class variable listing supported model IDs, for example gpt-5-mini and gpt-4o. To view all supported models go to the API reference or run:
from haystack.components.generators.chat import AzureOpenAIChatGenerator
print(AzureOpenAIChatGenerator.SUPPORTED_MODELS)
We will add this for other model providers in their respective ChatGenerator components step by step.
Added partial support for the image-text-to-text task in HuggingFaceLocalChatGenerator.
This allows the use of multimodal models like Qwen 3.5 or Ministral with text-only inputs. Complete multimodal support via Hugging Face Transformers might be addressed in the future.
Added async filter helpers to the InMemoryDocumentStore: update_by_filter_async(), count_documents_by_filter_async(), and count_unique_metadata_by_filter_async().
InMemoryDocumentStore: get_metadata_fields_info_async(), get_metadata_field_min_max_async(), and get_metadata_field_unique_values_async(). These rely on the store's thread-pool executor, consistent with the existing async method pattern._to_trace_dict method to ImageContent and FileContent dataclasses. When tracing is enabled, the large base64-encoded binary fields (base64_image and base64_data) are replaced with placeholder strings (e.g. "Base64 string (N characters)"), consistent with the behavior of ByteStream._to_trace_dict.T + T -> list[T], T + list[T] -> list[T], str + ChatMessage -> list[str], str + ChatMessage -> list[ChatMessage], and all other str <-> ChatMessage conversion variants. This enables pipeline patterns like joining a plain query string with a list of ChatMessage objects into a single list[ChatMessage] input without any extra components.Fixed an issue in ChatPromptBuilder where specially crafted template variables could be interpreted as structured content (e.g., images, tool calls) instead of plain text.
Template variables are now automatically sanitized during rendering, ensuring they are always treated as plain text.
DocumentCleaner. The warning for documents with None content used %{document_id} instead of {document_id}, preventing proper interpolation of the document ID.ToolInvoker._merge_tool_outputs silently appending None to list-typed state when a tool's outputs_to_state source key is absent from the tool result. This is a common scenario with PipelineTool wrapping a pipeline that has conditional branches where not all outputs are always produced even if defined in outputs_to_state. The mapping is now skipped entirely when the source key is not present in the result dict.InMemoryDocumentStore.write_documents that caused the BM25 average document length to be systematically underestimated.$defs/$ref in tool parameter schemas before sending them to the HuggingFace API. The HuggingFace API does not support JSON Schema $defs references, which are generated by Pydantic when tool parameters contain dataclass types. This fix inlines all $ref pointers and removes the $defs section from tool schemas in HuggingFaceAPIChatGenerator.bm25_tokenization_regex in InMemoryDocumentStore now uses r"(?u)\b\w+\b", including single-character words (e.g., "a", "C") in BM25 scoring. Previously, the regex r"(?u)\b\w\w+\b" excluded these tokens. This change may slightly alter retrieval results. To restore the old behavior, explicitly pass the previous regex when initializing the document store.@aayushbaluni, @anakin87, @bilgeyucel, @bogdankostic, @Br1an67, @ComeOnOliver, @davidsbatista, @jnMetaCode, @julian-risch, @Krishnachaitanyakc, @maxdswain, @pandego, @RMartinWhozfoxy, @satishkc7, @sjrl, @srini047, @SyedShahmeerAli12, @v-tan, @xr843
Nothing published for this version
Fixed an issue in ChatPromptBuilder where specially crafted template variables could be interpreted as structured content (e.g., images, tool calls) i
ChatPromptBuilder where specially crafted template variables could be interpreted as structured content (e.g., images, tool calls) instead of plain text. Template variables are now automatically sanitized during rendering, ensuring they are always treated as plain text.DocumentCleaner. The warning for documents with None content used %{document_id} instead of {document_id}, preventing proper interpolation of the document ID.Nothing published for this version
Agent now supports Jinja2 templating in system_prompt, enabling runtime parameter injection and conditional logic directly in system messages. This ma
Agent now supports Jinja2 templating in system_prompt, enabling runtime parameter injection and conditional logic directly in system messages. This makes it easier to adapt agent behavior dynamically (e.g. language, tone, time-aware responses) and reuse agents across contexts without redefining prompts
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=[weather_tool],
system_prompt="""{% message role='system' %}
You always respond in {{language}}.
{% endmessage %}""",
required_variables=["language"],
)
result = agent.run(
messages=[ChatMessage.from_user("What is the weather in London?")],
language="Italian" # required variable for the system prompt
)
print(result["last_message"].text)
# >> Il tempo a Londra è soleggiato.
LLMRanker for LLM-Based RerankingLLMRanker introduces LLM-powered reranking, treating relevance as a semantic reasoning task rather than similarity scoring. This can yield better results for complex or multi-step queries compared to cross-encoders. The component can also filter out irrelevant/duplicate documents entirely, helping provide higher-quality context in RAG pipelines and agent workflows while keeping context windows lean.
from haystack import Document
from haystack.components.rankers import LLMRanker
ranker = LLMRanker()
documents = [
Document(id="paris", content="Paris is the capital of France."),
Document(id="berlin", content="Berlin is the capital of Germany."),
]
result = ranker.run(query="capital of Germany", documents=documents)
print(result["documents"][0].id) # "berlin"
OpenAIChatGenerator, OpenAIResponsesChatGenerator, and AzureOpenAIResponsesChatGenerator now expose a SUPPORTED_MODELS class variable, giving you the list of models supported by each component.
from haystack.components.generators.chat import OpenAIChatGenerator
print(OpenAIChatGenerator.SUPPORTED_MODELS)
Note: We’ll roll out this pattern to other ChatGenerator components step by step. See issue #10627 for progress.
Added LLMRanker, a new ranker component that uses a ChatGenerator and PromptBuilder to rerank documents based on JSON-formatted LLM output. LLMRanker supports configurable prompts, optional custom chat generators, runtime top_k overrides, and serialization.
AzureOpenAIResponsesChatGenerator exposes a SUPPORTED_MODELS class variable listing supported model IDs, for example gpt-5-mini and gpt-4o. To view all supported models go to the [API reference](https://docs.haystack.deepset.ai/reference/generators-api#azureopenairesponseschatgenerator) or run:
from haystack.components.generators.chat import AzureOpenAIResponsesChatGenerator
print(AzureOpenAIResponsesChatGenerator.SUPPORTED_MODELS)
We now allow a component's whose input type is typed as a union of lists (e.g. list[str] | list[ChatMessage]) to allow multiple input connections. Previously we only supported bare lists (e.g. list[str]) or optional lists (e.g. list[str] | None) to allow multiple input connections. A common use case for this is using the AnswerBuilder component which has it's replies input typed as list[str] | list[ChatMessage].
The system_prompt initialization parameter of the Agent component now supports Jinja2 message template syntax. This allows you to define the template at initialization time and pass runtime variables when calling the run method. This can be useful to inject dynamic values (such as the current time) or to add conditional instructions.
Example usage:
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import tool
@tool
def weather(location: str) -> str:
return f"The weather in {location} is sunny."
agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=[weather],
system_prompt="""{% message role='system' %}
You always respond in {{language}}.
{% endmessage %}""",
required_variables=["language"],
)
messages = [ChatMessage.from_user("What is the weather in London?")]
result = agent.run(messages=messages, language="Italian")
print(result["last_message"].text)
# >> Il tempo a Londra è soleggiato.
OpenAIChatGenerator exposes a SUPPORTED_MODELS class variable listing supported model IDs, for example gpt-5-mini and gpt-4o. To view all supported models go to the [API reference](https://docs.haystack.deepset.ai/reference/generators-api#openaichatgenerator) or run:
from haystack.components.generators.chat import OpenAIChatGenerator
print(OpenAIChatGenerator.SUPPORTED_MODELS)
SearchableToolset now supports customizing the bootstrap search tool's name, description, and parameter descriptions via three new optional __init__ parameters: search_tool_name, search_tool_description, and search_tool_parameters_description. This allows users to tune the LLM-facing metadata to work better with different models.
Example usage:
from haystack.tools import SearchableToolset
toolset = SearchableToolset(
catalog=my_tools,
search_tool_name="find_tools",
search_tool_description="Find tools by keyword. Pass 1-3 words, not sentences.",
search_tool_parameters_description={
"tool_keywords": "Single words only, e.g. 'hotel booking'.",
},
)
Add support for python 3.14 to Haystack. Haystack already mostly supported python 3.14. Only minor changes were needed in regards to our type serialization and type checking when handling bare Union types.
Make the runtime parameter messages to Agent messages optional since it is possible to execute the agent with only providing a user_prompt.
Improve performance of HuggingFaceAPIDocumentEmbedder.run_async by requesting embedding inference concurrently. This can be controlled using the new concurrency_limit parameter.
Removed redundant deepcopy operations from Pipeline and AsyncPipeline execution. Component outputs are no longer deepcopied when collecting pipeline results, as inputs are already deepcopied before each component executes, preventing unintended mutations. Component inputs and outputs are also no longer deepcopied before being stored in tracing spans. These changes improve pipeline execution performance, especially when large objects (e.g., lists of Documents) flow between components and when include_outputs_from is used.
Reduced unnecessary deepcopies in Agent for improved performance. Replaced deepcopy of state_schema with a shallow dict copy since only top-level keys are modified, and removed deepcopy of agent_inputs for span tags since the dict is freshly created and only used for tracing.
Enable async embedding-based splitting with a new run_async method on EmbeddingBasedDocumentSplitter.
Added gpt-5.4 to OpenAIChatGenerator's list of supported models.
Added runtime validation of component output keys in Pipeline and AsyncPipeline. When a component returns keys that were not declared in its @component.output_types, the pipeline now logs a warning identifying the misconfigured component. This helps diagnose issues where a component returns unexpected keys, which previously caused a confusing "Pipeline Blocked" error pointing to an unexpected (downstream) component.
Fixed Agent.run_async to mirror Agent.run crash handling for internal chat_generator and tool_invoker failures. Async runs now wrap internal PipelineRuntimeError exceptions with Agent context and attach pipeline snapshots so standalone async failures can be debugged and resumed consistently.
Fixed ToolBreakpoint validation in Agent.run and Agent.run_async to validate against tools selected for the current run. This allows breakpoints for runtime tool overrides to work correctly.
Replaced in-place dataclass attribute mutation with dataclasses.replace() across multiple components to prevent unintended side-effects when the same dataclass instance is shared across pipeline branches.
Affected components and dataclasses:
ChatPromptBuilder and DynamicChatPromptBuilder: ChatMessage._contentHuggingFaceLocalChatGenerator: ChatMessage._contentHuggingFaceTEIRanker: Document.scoreMetaFieldRanker: Document.scoreSentenceTransformersSimilarityRanker: Document.scoreTransformersSimilarityRanker: Document.scoreExtractiveReader: ExtractedAnswer.queryInMemoryDocumentStore: Document.embeddingUpdate Pipeline.inputs() to return any variadic inputs as not mandatory if they already have a connection. Removed the utility functions describe_pipeline_inputs and describe_pipeline_inputs_as_string from haystack/core/pipeline/descriptions.py since they were not used and not referenced in the documentation. Use the Pipeline.inputs() method to inspect the inputs of a pipeline.
Fixed a bug in the pipeline scheduling logic where a component with all-optional inputs (e.g. Agent after making messages optional) could be scheduled ahead of a variadic joiner (e.g. ListJoiner) that was still waiting on inputs. The fix updates the tiebreaking logic in _tiebreak_waiting_components so that variadic joiners and components with all-optional inputs are treated at the same priority level, with topological order determining which waiting component runs first.
Fix ConditionalRouter incorrectly validating a plain str as list[str]. Since str is a Sequence, it previously passed the Sequence type check. Now str and bytes values are explicitly rejected when the expected type is a generic Sequence like list[str].
Use TypeVar instead of type as the type hint for cls in _warn_on_inplace_mutation. Using type would type the function as (type) -> type, losing information about which class was passed in. With (cls: T) -> T, the type checker understands that the specific class passed in is returned unchanged, rather than an anonymous type.
Improved the warning message emitted when the pipeline appears to be blocked. The message now lists all potentially affected components along with their types, and clarifies that some components may be intentionally inactive due to conditional branching. Previously the error message only listed one of the potentially blocking components and sometimes erroneously identified the wrong component as the blocker.
@Aftabbs, @agnieszka-m, @anakin87, @B-Step62, @bogdankostic, @Br1an67, @davidsbatista, @it-education-md, @jnMetaCode, @julian-risch, @kacperlukawski, @marc-mrt, @maxdswain, @rob-9, @sjrl, @travellingsoldier85, @Waqar53
Nothing published for this version
Reverts the change that made Agent messages optional as it caused issues with pipeline execution. As a consequence, the LLM component now defaults to
Nothing published for this version
Auto variadic sockets now also support Optional[list[...]] input types, in addition to plain list[...].
Optional[list[...]] input types, in addition to plain list[...].Optional[list[...]] (e.g. list[ChatMessage] | None). Previously, connecting two list[ChatMessage] outputs to Agent.messages would fail after its type was updated from list[ChatMessage] to list[ChatMessage] | None.Nothing published for this version
Removed the deprecated PipelineTemplate and PredefinedPipeline classes, along with the Pipeline.from_template() method. Users should migrate to Pipeli…
SearchableToolsetFor applications with large tool catalogs, we’ve added the SearchableToolset. Instead of exposing all tools upfront, agents start with a single search_tools function and dynamically discover relevant tools using BM25-based keyword search.
This is particularly useful when connecting MCP servers via MCPToolset, where many tools may be available. By combining the two, agents can load only the tools they actually need at runtime, reducing context usage and improving tool selection.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import Tool, SearchableToolset
# Create a catalog of tools
catalog = [
Tool(name="get_weather", description="Get weather for a city", ...),
Tool(name="search_web", description="Search the web", ...),
# ... 100s more tools
]
toolset = SearchableToolset(catalog=catalog)
agent = Agent(chat_generator=OpenAIChatGenerator(), tools=toolset)
# The agent is initially provided only with the search_tools tool and will use it to find relevant tools.
result = agent.run(messages=[ChatMessage.from_user("What's the weather in Milan?")])
Agents now natively support Jinja2-templated user prompts. By defining a user_prompt and required_variables during initialization or at runtime, you can easily invoke the Agent with dynamic variables without having to manually build ChatMessage objects for every invocation. Plus, you can seamlessly append rendered prompts directly to prior conversation messages.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=tools,
system_prompt="You are a helpful translation assistant.",
user_prompt="""{% message role="user"%}
Now summarize the conversation in {{ language }}.
{% endmessage %}""",
required_variables=["language"],
)
result = agent.run(
messages=[
ChatMessage.from_user("What are the main benefits of renewable energy?"),
ChatMessage.from_assistant("Renewable energy reduces greenhouse gas emissions, decreases dependence on fossil fuels, and can lower long-term energy costs."),
],
language="Spanish",
)
Removed the deprecated PipelineTemplate and PredefinedPipeline classes, along with the Pipeline.from_template() method. Users should migrate to Pipeline YAML files for similar functionality. See the [Serialization documentation](https://docs.haystack.deepset.ai/docs/serialization) for details on using YAML-based pipeline definitions.
Default Hugging Face pipeline task updated to ``text-generation``
The default task used by HuggingFaceLocalGenerator has been changed from text2text-generation to text-generation and the default model has been changed from "google/flan-t5-base" to "Qwen/Qwen3-0.6B".
In transformers v5+, text2text-generation is no longer available as a valid pipeline task (see: https://github.com/huggingface/transformers/pull/43256). While parts of the implementation still exist internally, it is no longer supported as a straightforward pipeline option.
How to know if you are affected
transformers>=5.0.0.task="text2text-generation" in HuggingFaceLocalGenerator or HuggingFaceLocalChatGenerator.How to handle this change
task="text2text-generation" with task="text-generation".text-generation pipeline (for example, causal language models).transformers<5.text2text-generation is now considered deprecated in Haystack and may be removed in a future release.Added link_format parameter to PPTXToDocument and XLSXToDocument converters, allowing extraction of hyperlink addresses from PPTX and XLSX files.
Supported formats:
"markdown": [text](url)"plain": text (url)"none" (default): Only text is extracted, link addresses are ignored.This follows the same pattern already available in DOCXToDocument.
Added a new LLM component (haystack.components.generators.chat.LLM) that provides a simplified interface for text generation powered by a large language model. The LLM component is a streamlined version of the Agent that focuses solely on single-turn text generation without tool usage. It supports system prompts, templated user prompts with required variables, streaming callbacks, and both synchronous (run) and asynchronous (run_async) execution.
Usage example:
from haystack.components.generators.chat import LLM
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
llm = LLM(
chat_generator=OpenAIChatGenerator(),
system_prompt="You are a helpful translation assistant.",
user_prompt="""{% message role="user"%}
Summarize the following document: {{ document }}
{% endmessage %}""",
required_variables=["document"],
)
result = llm.run(document="The weather is lovely today and the sun is shining. ")
print(result["last_message"].text)
Added SearchableToolset to haystack.tools module. This new toolset enables agents to dynamically discover tools from large catalogs using keyword-based (BM25) search. Instead of exposing all tools upfront (which can overwhelm LLMs with large tool definitions), agents start with a single search_tools function and progressively discover relevant tools as needed. For smaller catalogs, it operates in passthrough mode exposing all tools directly.
Key features include configurable search threshold for automatic passthrough mode and top-k result limiting.
Added user_prompt and required_variables parameters to the Agent component. You can now define a reusable Jinja2-templated user prompt at initialization or at runtime, so the Agent can be invoked with different inputs without manually constructing ChatMessage objects each time.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=tools,
system_prompt="You are a helpful translation assistant.",
user_prompt="""{% message role="user"%}
Translate the following document to {{ language }}: {{ document }}
{% endmessage %}""",
required_variables=["language", "document"],
)
result = agent.run(language="French", document="The weather is lovely today.")
When you combine messages with user_prompt, the rendered user prompt is appended to the provided messages. This is useful for passing prior conversation context alongside a new templated query.
Added the FileToFileContent component, which converts local files into FileContent objects. These can be embedded into ChatMessage to pass to an LLM.
Added document_comparison_field parameter to DocumentMRREvaluator, DocumentMAPEvaluator, and DocumentRecallEvaluator.
This allows users to compare documents using fields other than content, such as id or metadata keys (via meta.<key> syntax).
Previously, all three evaluators hardcoded doc.content for comparison, which did not work well when documents were chunked or when ground truth was identified by custom metadata fields.
Added support for transformers v5. Haystack remains fully compatible with transformers v4, but upgrading to v5 unlocks several benefits, including faster model loading times, improved model quantization support, faster inference for selected models, and other underlying improvements. Read more in the transformers v5 release blog post.
The LLMDocumentContentExtractor now extracts both content and metadata from image-based documents. When the LLM returns JSON, document_content fills the document body and other keys are merged into metadata; plain text is still used as content. The field content_extraction_error is no longer used and when an error occurs the field extraction_erroris added to metadata with the error message.
Improved the deserialization error message for pipeline components to be more actionable and human-readable. The component data dictionary is now pretty-printed as formatted JSON, and the underlying error that caused the failure is explicitly surfaced, making it easier to quickly diagnose deserialization issues.
EmbeddingBasedDocumentSplitter and MultiQueryEmbeddingRetriever now automatically invoke warm_up() when run() is called if they have not been warmed up yet.
Improved ComponentTool to correctly handle components whose run method parameters are declared as top-level Optional types such as list[ChatMessage] | None. The optional wrapper is now unwrapped before checking for a from_dict method on the underlying type. As a result, when a parameter is typed as list[ChatMessage] | None and receives a list of dictionaries, ComponentTool will automatically coerce the input into a list of ChatMessage objects using ChatMessage.from_dict. If the provided value is None, the parameter is preserved as None.
Haystack now emits a Warning when dataclass instances (e.g. Document, ChatMessage, StreamingChunk, ByteStream, SparseEmbedding) are mutated in place. Modifying shared instances can cause unexpected behavior in other parts of the pipeline. Use dataclasses.replace to safely create updated copies instead.
Instead of modifying attributes in place:
from haystack.dataclasses import Document
doc = Document(content="old text", meta={"key": "value"})
# Not recommended: can affect other parts of the pipeline
doc.content = "new text"
Use dataclasses.replace to create a new instance with the updated values:
from dataclasses import replace
from haystack.dataclasses import Document
doc = Document(content="old text", meta={"key": "value"})
# Recommended: creates a new Document with updated content
doc = replace(doc, content="new text")
Fixed an issue in OpenAIChatGenerator and OpenAIResponsesChatGenerator where passing a FileContent object without a filename would raise an error. A fallback filename is now automatically used instead.
Ensure type display works correctly for parameterized generics when tracing is enabled. Previously, the haystack.component.input_spec and haystack.component.output_spec tag would strip the arguments present within a container type (e.g. list[str] would become "list"). Now we properly keep the arguments representation (e.g. list[str] becomes "list[str]").
Previously, flexible pipeline connections were not robust when the sender component returned a list containing a union of types. In flexible pipeline connections, if the receiver type is str or ChatMessage, the first element of the list sent by the sender component is extracted and converted to the receiver type.
In the previous version, list[str | int] was considered compatible with str, which should not be the case. In fact, the sender component can legitimately return a list where the first element has type int, which the receiver cannot handle.
This is now fixed by ensuring that all possible element types of the sender list can be converted to the receiver type using the same conversion strategy.
Improved device handling when loading Hugging Face models in TransformersSimilarityRanker and ExtractiveReader.
hf_device_map is not always present anymore and is now only set when mixed-device loading is explicitly configured. The code has been updated to:
hf_device_map is available.device attribute when it is not.This prevents attribute errors and ensures compatibility across different transformers configurations.
Updated failing unit tests to align with recent mocking and transformers behavior changes.
PipelineRuntimeError raised by Agent now provide clearer ownership by explicitly surfacing the Agent as the failing pipeline component.
As a result:
component_name now resolves to the name of the Agent in the pipeline, instead of the underlying chat_generator or tool_invoker.component_type now resolves to haystack.components.agents.agent.Agent instead of the concrete generator class such as haystack.components.generators.chat.openai.OpenAIChatGenerator.Agent component.Example of new error message:
The following component failed to run:
Component name: 'agent'
Component type: 'Agent'
Error: The following component failed to run:
Component name: 'chat_generator'
Component type: 'OpenAIChatGenerator'
Error: Error code: 404 - {'error': {'message': 'The model ``gpt-4.2-mini`` does not exist or you do not have access to it.', 'type': 'invalid_request_error', 'param': None, 'code': 'model_not_found'}}
@agnieszka-m, @Amanbig, @anakin87, @bilgeyucel, @bogdankostic, @davidsbatista, @edwiniac, @julian-risch, @kacperlukawski, @marc-mrt, @OGuggenbuehl, @OiPunk, @sjrl, @vblagoje, @yaowubarbara
Nothing published for this version
Fixed a bug in flexible Pipeline connections that prevented automatic value conversion when the receiving component expects a Union type. For example,
ChatMessage to a receiver expecting list[str] | list[ChatMessage] should have worked but did not. The conversion strategy now correctly evaluates each branch of a Union receiver and picks the best match.Nothing published for this version
⚠️ Breaking change: MultiQueryEmbeddingRetriever and MultiQueryTextRetriever now also deduplicate by _id_ (not by content). Documents are only conside…
With the updated logic, Pipelines can now:
list[T] component outputs directly to a single list[T] input of the next component, simplifying pipeline definitions when multiple components produce compatible outputs. E.g., you can directly connect multiple converters to a writer component without a DocumentJoiner component in ingestion pipelines.ChatMessage and str types, enabling simpler connections between various components. E.g., you can easily connect an Agent component (which returns ChatMessage as last_message) to a text embedder component (which expects a str as query) without an OutputAdapter component.list[ChatMessage] to ChatMessage and list[str] to str by taking the first element of the list. Another supported conversion is list[ChatMessage] to str, enabling the connection between a chat generator (which returns list[ChatMessage] as messages) and a BM25 retriever (which expects a str as query).T can be connected to a component expecting type list[T].Together, these changes eliminate the need for OutputAdapter and some joiners (ListJoiner, DocumentJoiner) in many common setups such as query rewriting and hybrid search.
Rankers now deduplicate documents by id before ranking, preventing the same document from being scored multiple times in hybrid retrieval setups and eliminating the need for a DocumentJoiner after recent pipeline updates.
⚠️ Breaking change: MultiQueryEmbeddingRetriever and MultiQueryTextRetriever now also deduplicate by id (not by content). Documents are only considered duplicates if they share the same id.
You can now include files (e.g. PDFs) in ChatMessage objects using the new FileContent dataclass when working with OpenAI and Azure chat generators, including Responses API. Support for more providers is coming soon (see #10474)
from haystack.components.generators.chat.openai import OpenAIChatGenerator
from haystack.dataclasses.chat_message import ChatMessage
from haystack.dataclasses.file_content import FileContent
file_content = FileContent.from_url("https://arxiv.org/pdf/2309.08632")
chat_message = ChatMessage.from_user(content_parts=[file_content, "Summarize this paper in 100 words."])
llm = OpenAIChatGenerator(model="gpt-4.1-mini")
response = llm.run(messages=[chat_message])
Deduplication in Rankers is a breaking change for users who rely on keeping duplicate documents with the same user-defined id in the ranking output. This change only affects users with custom document ids who want duplicates preserved. To keep the previous behavior, ensure that your user-defined document ids are unique across retriever outputs.
Affected Rankers: HuggingFaceTEIRanker, LostInTheMiddleRanker, MetaFieldRanker, MetaFieldGroupingRanker, SentenceTransformersDiversityRanker, SentenceTransformersSimilarityRanker, TransformersSimilarityRanker.
Deduplication behavior in MultiQueryEmbeddingRetriever and MultiQueryTextRetriever has changed and may be breaking for users who relied on deduplication based on document content rather than document id.
Documents are now considered duplicates only if they share the same id. Document ids can be user-defined or are automatically generated as a hash of the document’s attributes (e.g. content, metadata, etc.).
This change affects setups where multiple documents have identical content but different ids (for example, due to differing metadata). To preserve the previous behavior, ensure that documents with identical content are assigned the same id across retriever outputs.
Removed the deprecated deserialize_document_store_in_init_params_inplace function. This function was deprecated in Haystack 2.23.0 and is no longer used.
Pipelines now natively support connecting multiple outputs directly to a single component input without requiring an explicit Joiner component. This only works when the connected outputs and inputs are of compatible list types, such as list[Document].
This simplifies pipeline definitions when multiple components produce compatible outputs. For example, multiple outputs from a FileTypeRouter can now be connected directly to a single converter or writer, without defining an intermediate ListJoiner or DocumentJoiner.
from haystack import Pipeline
from haystack.components.converters import HTMLToDocument, TextFileToDocument
from haystack.components.routers import FileTypeRouter
from haystack.components.writers import DocumentWriter
from haystack.dataclasses import ByteStream
from haystack.document_stores.in_memory import InMemoryDocumentStore
sources = [
ByteStream.from_string(text="Text file content", mime_type="text/plain", meta={"file_type": "txt"}),
ByteStream.from_string(
text="\n<html><body>Some content</body></html>\n", mime_type="text/html", meta={"file_type": "html"},
),
]
doc_store = InMemoryDocumentStore()
pipe = Pipeline()
pipe.add_component("router", FileTypeRouter(mime_types=["text/plain", "text/html"]))
pipe.add_component("txt_converter", TextFileToDocument())
pipe.add_component("html_converter", HTMLToDocument())
pipe.add_component("writer", DocumentWriter(doc_store))
pipe.connect("router.text/plain", "txt_converter.sources")
pipe.connect("router.text/html", "html_converter.sources")
# The DocumentWriter accepts documents from both converters without needing a DocumentJoiner
pipe.connect("txt_converter.documents", "writer.documents")
pipe.connect("html_converter.documents", "writer.documents")
result = pipe.run({"router": {"sources": sources}})
# result["writer"]["documents_written"] == 2
Pipelines now support connection and automatic conversion between ChatMessage and str types.
- When a str output is connected to a ChatMessage input, it is automatically converted to a user ChatMessage.
ChatMessage output is connected to a str input, its text attribute is automatically extracted. If text is None, an informative PipelineRuntimeError is raised.Pipelines now support list wrapping: a component returning type T can be connected to a component expecting type list[T]. The output will be wrapped automatically.
In addition, Pipelines support automatic conversion between list[T] and T, for str and ChatMessage types only. When converting from list[T] to T, the first element of the list is used. If the list is empty, an informative PipelineRuntimeError is raised.
With other recent changes, this makes pipelines more flexible and removes the need for explicit adapter components in many cases. For example, the following pipeline automatically converts a list[ChatMessage] produced by the LLM into a str expected by the retriever, which previously required an Output Adapter component.
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.dataclasses import Document
from haystack.components.retrievers import InMemoryBM25Retriever
from haystack import Pipeline
from haystack.components.builders import ChatPromptBuilder
from haystack.components.generators.chat import OpenAIChatGenerator
document_store = InMemoryDocumentStore()
documents = [
Document(content="Bob lives in Paris."),
Document(content="Alice lives in London."),
Document(content="Ivy lives in Melbourne."),
Document(content="Kate lives in Brisbane."),
Document(content="Liam lives in Adelaide."),
]
document_store.write_documents(documents)
template ="""{% message role="user" %}
Rewrite the following query to be used for keyword search.
{{ query }}
{% endmessage %}
"""
p = Pipeline()
p.add_component("prompt_builder", ChatPromptBuilder(template=template))
p.add_component("llm", OpenAIChatGenerator(model="gpt-4.1-mini"))
p.add_component("retriever", InMemoryBM25Retriever(document_store=document_store, top_k=3))
p.connect("prompt_builder", "llm")
p.connect("llm", "retriever")
query = """Someday I'd love to visit Brisbane, but for now I just want
to know the names of the people who live there."""
results = p.run(data={"prompt_builder": {"query": query}})
Introduced the MarkdownHeaderSplitter component:
#, ##, etc.), preserving header hierarchy as metadata.DocumentSplitter.Added support for Chat Messages that include files using the FileContent dataclass to OpenAIResponsesChatGenerator and AzureOpenAIResponsesChatGenerator.
Users can now pass files such as PDFs when using Haystack Chat Generators based on the Responses API.
User Chat Messages can now include files using the new FileContent dataclass.
Most API-based LLMs support file inputs such as PDFs (and, for some models, additional file types).
For now, this feature is implemented for OpenAIChatGenerator and AzureOpenAIChatGenerator, with support for more model providers coming soon.
For advanced PDF handling, such as reducing image size or selecting specific page ranges, we recommend using the PDFToImageContent component instead.
FileContent example:
from haystack.components.generators.chat.openai import OpenAIChatGenerator
from haystack.dataclasses.chat_message import ChatMessage
from haystack.dataclasses.file_content import FileContent
file_content = FileContent.from_url("https://arxiv.org/pdf/2309.08632")
chat_message = ChatMessage.from_user(content_parts=[file_content, "Summarize this paper in 100 words."])
llm = OpenAIChatGenerator(model="gpt-4.1-mini")
response = llm.run(messages=[chat_message])
Resolve postponed type annotations (from from __future__ import annotations) when creating component input sockets, so pipelines can correctly match compatible types. This fixes cases where connecting ChatPromptBuilder to FallbackChatGenerator failed because the generator’s annotations were interpreted as strings (for example 'list[ChatMessage]'), resulting in a PipelineConnectError due to mismatched socket types.
Agent components allow grouping multiple tools under a single confirmation strategy.
Here is an example of how three tools can be grouped under a BlockingConfirmationStrategy:
confirmation_strategies = {
("tool1", "tool2", "tool3"): BlockingConfirmationStrategy()
}
instead of previously needing
confirmation_strategies = {
"tool1": BlockingConfirmationStrategy(),
"tool2": BlockingConfirmationStrategy(),
"tool3": BlockingConfirmationStrategy()
}
Add run_async method to SearchApiWebSearch and SerperDevWebSearch
Add new DocumentStore standard tests for the following operations: delete_all_documents(), update_by_filter(), delete_by_filter()
The InMemoryDocumentStore now has three new operations delete_all_documents(), update_by_filter() and delete_by_filter()
Rankers now deduplicate documents by id before ranking, preventing identical documents from being scored multiple times in hybrid retrieval setups and keeping ranking outputs more consistent.
It also means you can connect multiple retriever outputs directly to a Ranker without inserting a DocumentJoiner just to avoid duplicates. For example:
from haystack import Pipeline
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.rankers import TransformersSimilarityRanker
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever, InMemoryEmbeddingRetriever
from haystack.document_stores.in_memory import InMemoryDocumentStore
document_store = InMemoryDocumentStore()
text_embedder = SentenceTransformersTextEmbedder(model="BAAI/bge-small-en-v1.5")
embedding_retriever = InMemoryEmbeddingRetriever(document_store)
bm25_retriever = InMemoryBM25Retriever(document_store)
ranker = TransformersSimilarityRanker(model="BAAI/bge-reranker-base")
hybrid_retrieval = Pipeline()
hybrid_retrieval.add_component("text_embedder", text_embedder)
hybrid_retrieval.add_component("embedding_retriever", embedding_retriever)
hybrid_retrieval.add_component("bm25_retriever", bm25_retriever)
hybrid_retrieval.add_component("ranker", ranker)
hybrid_retrieval.connect("text_embedder", "embedding_retriever")
hybrid_retrieval.connect("embedding_retriever", "ranker")
hybrid_retrieval.connect("bm25_retriever", "ranker")
query = "apnea in infants"
result = hybrid_retrieval.run(
{"text_embedder": {"text": query}, "bm25_retriever": {"query": query}, "ranker": {"query": query}}
)
Add strip_whitespaces and replace_regexes parameters to DocumentCleaner component.
The strip_whitespaces parameter removes leading and trailing whitespace from document content using Python's str.strip()method. Unlike remove_extra_whitespaces, this only affects the beginning and end of the text, preserving internal whitespace which is useful for maintaining markdown formatting.
The replace_regexes parameter accepts a dictionary mapping regex patterns to replacement strings, allowing custom text transformations. For example, {r'\\n\\n+': '\\n'} replaces multiple consecutive newlines with a single newline. This is applied after remove_regex and provides more flexibility than simple pattern removal.
Example usage:
from haystack.components.preprocessors import DocumentCleaner
from haystack.dataclasses import Document
cleaner = DocumentCleaner(
strip_whitespaces=True,
replace_regexes={r'\n\n+': '\n'}
)
doc = Document(content=" \n\nHello World\n\n\n ")
result = cleaner.run(documents=[doc])
# Result: "Hello World\n"
The MultiQueryEmbeddingRetriever and MultiQueryTextRetriever now deduplicate documents by id instead of by content, preventing identical documents from being returned multiple times.
PipelineTemplate, PredefinedPipeline and its options (like PredefinedPipeline.CHAT_WITH_WEBSITE). These templates will be removed in Haystack 2.25. Users should switch to using Pipeline YAML files.Pipeline and AsyncPipeline deep-copies component inputs before execution so mutable outputs (e.g., Document dataclasses) shared across multiple downstream components don't get mutated by reference. This prevents side effects where one component's in-place modifications could unexpectedly affect other branches in the pipeline.@agnieszka-m, @Amanbig, @anakin87, @bilgeyucel, @Bobholamovic, @bogdankostic, @davidsbatista, @julian-risch, @kacperlukawski, @maxdswain, @OGuggenbuehl, @OiPunk, @sjrl, @srini047, @VedantMadane
Nothing published for this version
Nothing published for this version
deserialize_document_store_in_init_params_inplace is deprecated and will be removed in Haystack version 2.24. It is no longer used internally and shou…
Agents can now pause for human confirmation before executing tools. You can define confirmation behavior per tool: always ask, ask only on first use, or never ask and fully customize the confirmation UI. This makes it much easier to build safer, more transparent agent workflows, especially when tools trigger side effects or access sensitive data.
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-4.1"),
tools=[balance_tool, addition_tool, phone_tool],
system_prompt="You are a helpful financial assistant. Use the provided tool to get bank balances when needed.",
confirmation_strategies={
balance_tool.name: BlockingConfirmationStrategy(
confirmation_policy=AlwaysAskPolicy(), confirmation_ui=RichConsoleUI(console=cons),
),
phone_tool.name: BlockingConfirmationStrategy(
confirmation_policy=AskOncePolicy(), confirmation_ui=SimpleConsoleUI(),
),
addition_tool.name: BlockingConfirmationStrategy(
confirmation_policy=NeverAskPolicy(), confirmation_ui=SimpleConsoleUI(),
)
},
)
For a detailed walkthrough of confirmation strategies and UI customization, see Tutorial: Human-in-the-Loop with Haystack Agents
Tool classes can now return images alongside text, enabling richer agent responses.
ToolCallResult.result supports lists of TextContent and ImageContent, allowing agents to retrieve, pass around, and describe images when used with providers that support it (e.g. OpenAIResponsesChatGenerator, AnthropicChatGenerator and more).
This unlocks new use cases like image-based tool outputs like custom retrievals, visual search and inspection through MCP tools returning base64 strings, and multimodal agent reasoning.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.tools import ComponentTool
from haystack.dataclasses import ChatMessage, ImageContent
from haystack import component
@component
class ImageRetriever:
@component.output_types(images=list[ImageContent])
def run(self):
return {"images":[ImageContent.from_file_path("/content/image.jpg")]}
image_retriever_tool = ComponentTool(
component=ImageRetriever(), outputs_to_string={"raw_result":True, "source":"images"}
)
agent = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5-nano"),
system_prompt="You are an Agent that can retrieve images and describe them.",
tools=[image_retriever_tool],
)
user_message = ChatMessage.from_user("Retrieve the image and describe it. Tell me if you cannot see the image")
result = agent.run(messages=[user_message])
print(result["last_message"].text)
Custom components now serialize and deserialize automatically in most cases via component_from_dict() and component_to_dict(), even when they contain complex attributes such as DocumentStore, Secret,ComponentDevice, or any object that implements to_dict()/from_dict().
This means most custom components no longer need to implement to_dict()/from_dict() themselves. Pipeline snapshots and YAML definitions are easier to create and restore, and custom components are more portable and less error-prone.
HAYSTACK_PIPELINE_SNAPSHOT_SAVE_ENABLED=true.pipeline_outputs format. Pipeline snapshots created before Haystack 2.22.0 that contain pipeline_outputs without the serialization_schema and serialized_data structure are no longer supported. Users should recreate their pipeline snapshots with the current Haystack version before upgrading to 2.23.0.return_empty_on_no_match parameter has been fully removed from the RegexTextExtractor component. In Haystack 2.22.0, this parameter was ignored. Starting with Haystack 2.23.0, passing this parameter during component initialization will raise an error. During pipeline deserialization, the parameter is ignored to avoid breaking existing pipelines.Added a snapshot_callback parameter to Pipeline.run() that allows users to customize how pipeline snapshots are handled. When a callback is provided, it is invoked instead of the default file-saving behavior whenever a snapshot is created (e.g., during breakpoints or error handling). This enables use cases like saving snapshots to a database, sending them to a remote service, or implementing custom logging. If no callback is provided, the default behavior of saving to a JSON file remains unchanged.
component_from_dict() and component_to_dict() now work with custom components out of the box also if the component has a ComponentDevice as an attribute. Users no longer need to explicitly define to_dict() and from_dict() methods in their custom components to call ComponentDevice.from_dict() or device.to_dict(). component_from_dict() and component_to_dict() now handle this automatically.
component_from_dict/component_to_dict now work out of the box with custom components that have an object as init parameter as long as the object defines to_dict/from_dict methods. Users no longer need to explicitly define to_dict/from_dict methods in their custom components in such cases. For example, a custom retriever, which has a DocumentStore as an init parameter, does not need explicitly defined to_dict/from_dict methods. component_from_dict/component_to_dict now handle such cases automatically.
component_from_dict() and component_to_dict() now work with custom components out of the box also if the component has a Secret as an attribute. Users no longer need to explicitly define to_dict() and from_dict() methods in their custom components to call deserialize_secrets_inplace() or api_key.to_dict(). component_from_dict() and component_to_dict() now handle this automatically.
Added HAYSTACK_PIPELINE_SNAPSHOT_SAVE_ENABLED environment variable. When set to "true" or "1", pipeline snapshots are saved to disk. Disabled by default. Note: Custom snapshot_callback functions are still invoked regardless of this setting.
Expanded the ToolCallResult.result field to accept not only strings but also lists of TextContent and ImageContent objects. This enables tools to return images for providers that support this capability. This feature is already available when using OpenAIResponsesChatGenerator, and support for additional providers will be added soon. The Chat Completions API does not support this functionality, so the classic OpenAIChatGenerator cannot be used with it.
The outputs_to_string parameter of the Tool class now supports returning raw results without string conversion using the raw_result key. This is intended for tools that return images. ComponentTool and PipelineTool also support this feature.
Added haystack.component.fully_qualified_type field to component tracing output. This new field provides the full module path and class name (e.g., haystack.components.generators.chat.openai.OpenAIChatGenerator) alongside the existing haystack.component.type field that only contains the class name. This enables dynamic component loading and better tooling integration.
In OpenAIChatGenerator, streaming now handles cases where a ChatCompletionChunk has a delta set to None in choices. This can occur with some OpenAI-compatible providers, and the component will now handle it gracefully.
Components no longer handle the (de-)serialization of ComponentDevice explicitly. Instead, the components rely on the behavior implemented in default_to_dict/default_from_dict.
Components no longer handle the (de-)serialization of init parameter objects explicitly if the objects define to_dict/from_dict themselves. Instead, the components rely on the behavior implemented in default_to_dict/default_from_dict.
Components no longer handle the (de-)serialization of Secrets explicitly. Instead, the components rely on the behavior implemented in default_to_dict/default_from_dict.
Support for flattened generation_kwargs in OpenAIResponsesChatGenerator
The OpenAIResponsesChatGenerator component now supports flattened generation keyword arguments, allowing users to specify reasoning parameters directly without nesting them. This enhancement simplifies the configuration process and improves usability.
Example:
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
generator = OpenAIResponsesChatGenerator(
model="gpt-4",
generation_kwargs={
"reasoning_effort": "low",
"reasoning_summary": "auto"
}
)
Support for flattened verbosity in generation_kwargs of OpenAIResponsesChatGenerator
The OpenAIResponsesChatGenerator component now supports flattened verbosity generation keyword arguments, allowing users to specify verbositydirectly without nesting them in text. This enhancement simplifies the configuration process and improves usability.
Example:
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
generator = OpenAIResponsesChatGenerator(
model="gpt-4",
generation_kwargs={
"verbosity": "low",
}
)
Added the outputs_to_string parameter to create_tool_from_function and the @tool decorator to provide additional customization options for these convenience constructors.
deserialize_document_store_in_init_params_inplace is deprecated and will be removed in Haystack version 2.24. It is no longer used internally and should not be used in new code. The deserialization of DocumentStores is handled automatically now by default_from_dict.ComponentTool, create_tool_from_function, and the @tool decorator failing to create a tool schema when Callable type parameters are present (such as snapshot_callback). This enables using Agent as a ComponentTool without raising SchemaGenerationError.OpenAIResponsesChatGenerator where empty reasoning items were discarded during streaming. This caused subsequent requests to the OpenAI Responses API to fail when the message history was sent back, as the API requires every tool call to be preceded by its associated reasoning item. These items are now correctly preserved in the ChatMessage history, even when the summary text is empty.SASEvaluator to work when using numpy>=2.4 by manually squeezing a PyTorch tensor to the correct dimension.dict[str, Any] | None to dict[str, Any]. Instead we now keep the None in the final type for the super component.create instead of parse method of the OpenAI Responses python client when text is set but no text_format is set. parse requires a type from text_format to be set, to actually parse the response into that type. With text set but no text_format, we just want to create a normal response without parsing.Tool serialization. Previously, when outputs_to_string used the multiple-output format, handlers were not serialized correctly.@agnieszka-m, @anakin87, @bilgeyucel, @chenopis, @GunaPalanivel, @julian-risch, @majiayu000, @mpangrazzi, @sjrl, @tstadel
Nothing published for this version
Introducing the new EmbeddingBasedDocumentSplitter, a component that takes an embedder and splits documents based on semantic similarity rather than f
Introducing the new EmbeddingBasedDocumentSplitter, a component that takes an embedder and splits documents based on semantic similarity rather than fixed sizes or rules.
from haystack.components.embedders import SentenceTransformersDocumentEmbedder
from haystack.components.preprocessors import EmbeddingBasedDocumentSplitter
# Initialize an embedder to calculate semantic similarities
embedder = SentenceTransformersDocumentEmbedder()
# Configure the splitter with parameters that control splitting behavior
splitter = EmbeddingBasedDocumentSplitter(
document_embedder=embedder,
sentences_per_group=2, # Group 2 sentences before calculating embeddings
percentile=0.95, # Split when cosine distance exceeds 95th percentile
min_length=50, # Merge splits shorter than 50 characters
max_length=1000 # Further split chunks longer than 1000 characters
)
result = splitter.run(documents=[doc])
warm_up Runs Automatically on First UseComponents that define awarm_up method now run it automatically on first execution, removing the need for manual calls and preventing errors in standalone usage.
from haystack.components.embedders import SentenceTransformersTextEmbedder
text_embedder = SentenceTransformersTextEmbedder()
# text_embedder.warm_up() # ❌ Don't need this step anymore
print(text_embedder.run("I love pizza!"))
## {'embedding': [-0.07804739475250244, 0.1498992145061493,, ...]}
outputs_to_stringTools can now expose multiple string outputs via the new outputs_to_string configuration, giving you fine-grained control over how tool results are surfaced to the LLM, without changing the underlying tool logic.
def format_documents(documents):
return "\n".join(f"{i+1}. Document: {doc.content}" for i, doc in enumerate(documents))
def format_summary(metadata):
return f"Found {metadata['count']} results"
tool = Tool(
name="search",
description="Search for documents",
parameters={...},
function=search_func, # Returns {"documents": [Document(...)], "metadata": {"count": 5}, "debug_info": {...}}
outputs_to_string={
"formatted_docs": {"source": "documents", "handler": format_documents},
"summary": {"source": "metadata", "handler": format_summary}
# Note: "debug_info" is not included, so it won't be converted to a string
}
)
# After the tool invocation, the tool result includes:
# {
# "formatted_docs": "1. Document Title\n Content...\n2. ...",
# "summary": "Found 5 results"
# }
Haystack now requires Python 3.10 or later, as Python 3.9 reached End of Life (EOL) in October 2025.
HuggingFaceLocalChatGenerator now uses Qwen/Qwen3-0.6B as the default model, replacing the previous default.The parameters query_suffix and document_suffix have been added to SentenceTransformersSimilarityRanker to support the Qwen3 reranker model family.
Here is an example of how to use these new parameters to use the Qwen3-Reranker-0.6B:
from haystack import Document
from haystack.components.rankers.sentence_transformers_similarity import SentenceTransformersSimilarityRanker
ranker = SentenceTransformersSimilarityRanker(
model="tomaarsen/Qwen3-Reranker-0.6B-seq-cls",
query_prefix='<|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be "yes" or "no".<|im_end|>\n<|im_start|>user\n<Instruct>: Given a web search query, retrieve relevant passages that answer the query\n<Query>: ',
query_suffix="\n",
document_prefix="<Document>: ",
document_suffix="<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
)
result = ranker.run(
query="Which planet is known as the Red Planet?",
documents=[
Document(content="Venus is often called Earth's twin because of its similar size and proximity."),
Document(content="Mars, known for its reddish appearance, is often referred to as the Red Planet."),
Document(content="Jupiter, the largest planet in our solar system, has a prominent red spot."),
Document(content="Saturn, famous for its rings, is sometimes mistaken for the Red Planet."),
],
)
print(result)
NOTE: This only works with the Qwen3 reranker models that use the sequence classification architecture. For example, you can find some on tomaarsen's Hugging Face profile.
Added reasoning content support to HuggingFaceAPIChatGenerator. The component now extracts reasoning content from models that support chain-of-thought reasoning (e.g., DeepSeek R1). Both streaming and non-streaming modes are supported. Access via reply.reasoning.reasoning_text.
When an Agent runs as part of a Pipeline, the agent's tracing span now uses the component span as its parent. This enables proper nested trace visualization in tracing tools like Datadog, Braintrust, or OpenTelemetry backends.
The _handle_async_stream_response() method in OpenAIChatGenerator now handles asyncio.CancelledError exceptions. When a streaming task is cancelled mid-stream, the async for loop gracefully closes the stream using asyncio.shield() to ensure the cleanup operation completes even during cancellation.
A new enable_thinking parameter has been added to enable thinking mode in chat templates for thinking-capable models, allowing them to generate intermediate reasoning steps before producing final responses.
Add support for PEP 604 type syntax. This means that when defining types in components, you can use X | Y instead of Union[X, Y] and X | None instead of Optional[X]. The codebase has been migrated to the new syntax, but both syntaxes are fully supported.
Support Multiple Tool String Outputs
Added support for tools to define multiple string outputs using the outputs_to_string configuration. This allows users to specify how different parts of a tool's output should be converted to strings, enhancing flexibility in handling tool results.
ToolInvoker to handle multiple output configurations.Tool to validate and store multiple output configurations.This enables tools to provide rich, varied context to language models or downstream components without requiring multiple tool calls, while keeping full control over which outputs are stringified.
Added validation for inputs_from_state and outputs_to_state parameters in the Tool class. Tools now validate at construction time that state mappings reference valid tool parameters and outputs, catching configuration errors early instead of at runtime. The validation uses function introspection and JSON schema to ensure parameter names exist, and subclasses like ComponentTool validate against component input/output sockets.
ConditionalRouter, ChatPromptBuilder, PromptBuilder and OutputAdapter by properly skipping variables that are assigned within the template. Previously under specific scenarios variables assigned within a template would falsely be picked up as input variables to the component. For more information you can check out the parent issue in the Jinja2 library here: https://github.com/pallets/jinja/issues/2069NamedEntityExtractor when pipeline_kwargs is stored in the deserialization dict with the value of None.limits parameter to an httpx.Limits object to avoid AttributeError.ValueError when an async function is passed to the Tool class. Async functions are not supported as tools. This change provides a clear error message instead of silent failures where coroutines are never awaited.return_empty_on_no_match parameter has been removed from the RegexTextExtractor component. This component now always returns a dictionary with the key "captured_text"; the value can be an empty string if no match is found or the captured text. Currently, the return_empty_on_no_match parameter is ignored. Starting from Haystack 2.23.0, initializing the component with this parameter will raise an error.@anakin87, @ArzelaAscoIi, @bilgeyucel, @Bobholamovic, @davidsbatista, @dfokina, @GunaPalanivel, @majiayu000, @OliverZhangA, @sjrl, @TaMaN2031A, @tommasocerruti, @tstadel, @vblagoje, @YassineGabsi
Nothing published for this version
This release introduces three new components that significantly boost retrieval recall in RAG systems by expanding the user query and retrieving docum
This release introduces three new components that significantly boost retrieval recall in RAG systems by expanding the user query and retrieving documents across multiple reformulations:
QueryExpander generates semantically similar variations of a user query to broaden search coverage.MultiQueryTextRetriever runs multiple queries in parallel using a text-based retriever (e.g., BM25) and merges results by score.MultiQueryEmbeddingRetriever performs the same multi-query retrieval flow using embeddings, enabling richer semantic recall.Used together, these components create a multi-query retrieval pipeline that improves recall especially when queries are short or ambiguous.
from haystack import Pipeline
from haystack.components.query import QueryExpander
from haystack.components.retrievers import InMemoryBM25Retriever
from haystack.components.retrievers import MultiQueryTextRetriever
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.writers import DocumentWriter
from haystack import Document
from haystack.document_stores.types import DuplicatePolicy
# Sample documents
docs = [
Document(content="Renewable energy comes from natural sources like wind and sunlight."),
Document(content="Geothermal energy is heat from beneath the Earth's surface."),
Document(content="Hydropower generates electricity using flowing water."),
]
# Store documents
store = InMemoryDocumentStore()
writer = DocumentWriter(document_store=store, policy=DuplicatePolicy.SKIP)
writer.run(documents=docs)
# Components
expander = QueryExpander()
retriever = InMemoryBM25Retriever(document_store=store, top_k=1)
multi_retriever = MultiQueryTextRetriever(retriever=retriever)
# Expand and retrieve
expanded = expander.run(query="renewable energy")
results = multi_retriever.run(queries=expanded["queries"])
for doc in results["documents"]:
print(doc.content)
This pipeline expands "renewable energy" into multiple related queries, retrieves documents for each in parallel, and returns a richer set of relevant results — demonstrating how multi-query retrieval improves recall with minimal effort.
gpt-4o-mini to gpt-4.1-mini and the default API version from 2023-05-15 to 2024-12-01-preview for both AzureOpenAIGenerator and AzureOpenAIChatGenerator.QueryExpander, MultiQueryEmbeddingRetriever, MultiQueryTextRetriever. When used together, they allow a query to be expanded and each expansion is used to retrieve a potentially different set of documents.ByteStream and ImageContent, the payload sent to the tracing backend could become too large, hitting provider limits or causing performance degradation. We now replace these objects with string placeholders to avoid oversized payloads.Ensure request header keys are unique in link_content to prevent 400 Bad Request errors.
Some image providers return a 400 Bad Request when using ImageContent.from_url() because the User-Agent header appears multiple times with different casing (e.g., user-agent, User-Agent). This update normalizes header keys in a case-insensitive way, removes duplicates, and preserves only the last occurrence.
Fixed a bug where components explicitly listed in <span class="title-ref">include_outputs_from</span> would not appear in the pipeline results if they returned an empty dictionary. Now, any component specified in <span class="title-ref">include_outputs_from</span> will be included in the results regardless of whether its output is empty.
Fix the serialization and deserialization of pipeline_outputs in pipeline_snapshot and make it use the same schema as the rest of the pipeline state when running pipelines with breakpoints. The deserialization of the older format of pipeline_outputs without serialization schema is supported till Haystack 2.23.0.
Fixed ToolInvoker missing tools after warmup for lazy-initialized toolsets. The invoker now refreshes its tool registry post-warmup, ensuring replaced placeholders (e.g., MCPToolset with eager_connect=False) resolve to the actual tool names at invocation time.
@Amnah199, @anakin87, @davidsbatista, @dfokina, @mrchtr, @OscarPindaro, @schwartzadev, @sjrl, @TaMaN2031A, @vblagoje, @YassineGabsi, @ZeJ0hn
Nothing published for this version
Haystack now integrates the OpenAI's Responses API through the new OpenAIResponsesChatGenerator and AzureOpenAIResponsesChatGenerator components.
Haystack now integrates the OpenAI's Responses API through the new OpenAIResponsesChatGenerator and AzureOpenAIResponsesChatGenerator components.
This unlocks several advanced capabilities like:
Tool objects and Toolset instances.Example with reasoning and a web search tool:
from haystack.components.generators.chat import AzureOpenAIResponsesChatGenerator
from haystack.dataclasses import ChatMessage
# with `OpenAIResponsesChatGenerator`
chat_generator = OpenAIResponsesChatGenerator(
model="o3-mini",
generation_kwargs={"summary": "auto", "effort": "low"},
tools=[{"type": "web_search"}],
)
response = chat_generator.run(messages=[ChatMessage.from_user("What's a positive news story from today?")])
# with `AzureOpenAIResponsesChatGenerator`
chat_generator = AzureOpenAIResponsesChatGenerator(
azure_endpoint="https://example-resource.azure.openai.com/",
azure_deployment="gpt-5-mini",
generation_kwargs={"reasoning": {"effort": "low", "summary": "auto"}},
)
response = chat_generator.run(messages=[ChatMessage.from_user("What's Natural Language Processing?")])
print(response["replies"][0].text)
AzureOpenAIResponsesChatGenerator, a new component that integrates Azure OpenAI's Responses API into Haystack.OpenAIResponsesChatGenerator, a new component that integrates OpenAI's Responses API into Haystack.ChatMessage.meta for OpenAIChatGenerator and OpenAIResponsesChatGenerator.extra field to ToolCall and ToolCallDelta to store provider-specific information.PipelineSnapshots to work with pydantic BaseModels.SentenceWindowRetriever with a new run_async() method, allowing the retriever to be used in async pipelines and workflows.warm_up() method to all ChatGenerator components (OpenAIChatGenerator, AzureOpenAIChatGenerator, HuggingFaceAPIChatGenerator, HuggingFaceLocalChatGenerator, and FallbackChatGenerator) to properly initialize tools that require warm-up before pipeline execution. The warm_up() method is idempotent and follows the same pattern used in Agent and ToolInvoker components. This enables proper tool initialization in pipelines that use ChatGenerators with tools but without an Agent component.AnswerBuilder component now exposes a new parameter return_only_referenced_documents (default: True) that controls if only documents referenced in the replies are returned. Returned documents include two new fields in the meta dictionary:
source_index: the 1-based index of the document in the input listreferenced: a boolean value indicating if the document was referenced in the replies (only present if the reference_pattern parameter is provided).
These additions make it easier to display references and other sources within a RAG pipeline.generation_kwargs to the Agent component, allowing for more fine-grained control at run-time over chat generation.revision parameter to all Sentence Transformers embedder components (SentenceTransformersDocumentEmbedder, SentenceTransformersTextEmbedder, SentenceTransformersSparseDocumentEmbedder, and SentenceTransformersSparseTextEmbedder) to allow users to specify a specific model revision/version from the Hugging Face Hub. This enables pinning to a particular model version for reproducibility and stability.Agent, LLMMetadataExtractor, LLMMessagesRouter, and LLMDocumentContentExtractor to automatically call self.warm_up() at runtime if they have not been warmed up yet. This ensures that the components are ready for use without requiring an explicit warm-up call. This differs from previous behavior where warm-up had to be manually invoked before use, otherwise a RuntimeError was raised.DatadogTracer by using the official ddtrace.tracer.get_log_correlation_context() method.Toolset.warm_up() method now warms up all tools by default, while subclasses can override it to customize initialization (e.g., setting up shared resources instead of warming individual tools). The warm_up_tools() utility function has been simplified to delegate to Toolset.warm_up().Fixed deserialization of state schema when it is None in Agent.from_dict.
Fixed a bug where components explicitly listed in include_outputs_from would not appear in the pipeline results if they returned an empty dictionary. Now, any component specified in include_outputs_from will be included in the results regardless of whether its output is empty.
Fixed type compatibility issue where passing list[Tool] to components with a tools parameter (such as ToolInvoker) caused static type checker errors.
In version 2.19, the ToolsType was changed to Union[list[Union[Tool, Toolset]], Toolset] to support mixing Tools and Toolsets. However, due to Python's list invariance, list[Tool] was no longer considered compatible with list[Union[Tool, Toolset]], breaking type checking for the common pattern of passing a list of Tool objects.
The fix explicitly lists all valid type combinations in ToolsType: Union[list[Tool], list[Toolset], list[Union[Tool, Toolset]], Toolset]. This preserves backward compatibility for existing code while still supporting the new functionality of mixing Tools and Toolsets.
Users who encountered type errors like "Argument of type 'list[Tool]' cannot be assigned to parameter 'tools'" should no longer see these errors after upgrading. No code changes are required on the user side.
When creating a pipeline snapshot, we now ensure use of _deepcopy_with_exceptions when copying component inputs to avoid deep copies of items like components and tools since they often contain attributes that are not deep-copyable.
For example, the LinkContentFetcher has httpx.Client as an attribute, which throws an error if deep-copied.
@Amnah199, @anakin87, @cmnemoi, @davidsbatista, @dfokina, @HamidOna, @Hansehart, @jdb78, @mrchtr, @sjrl, @swapniel99, @TaMaN2031A, @tstadel, @vblagoje
Nothing published for this version
Nothing published for this version
Introduced `FallbackChatGenerator`, a resilient chat generator that runs multiple LLMs sequentially and automatically falls back when one fails. It tr
FallbackChatGeneratorIntroduced FallbackChatGenerator, a resilient chat generator that runs multiple LLMs sequentially and automatically falls back when one fails. It tries each generator in order until one succeeds, handling errors like timeouts, rate limits, or server issues. Ideal for building robust, production-grade chat systems that stay responsive across providers.
from haystack.dataclasses import ChatMessage
from haystack_integrations.components.generators.google_genai import GoogleGenAIChatGenerator
from haystack_integrations.components.generators.anthropic import AnthropicChatGenerator
from haystack.components.generators.chat.openai import OpenAIChatGenerator
from haystack.components.generators.chat.fallback import FallbackChatGenerator
anthropic_generator = AnthropicChatGenerator(model="claude-sonnet-4-5", timeout=1) # force failure with low timeout
google_generator = GoogleGenAIChatGenerator(model="gemini-2.5-flashy") # force failure with typo in model name
openai_generator = OpenAIChatGenerator(model="gpt-4o-mini") # success
chat_generator = FallbackChatGenerator(chat_generators=[anthropic_generator, google_generator, openai_generator])
response = chat_generator.run(messages=[ChatMessage.from_user("What is the plot twist in Shawshank Redemption?")])
print("Successful ChatGenerator: ", response["meta"]["successful_chat_generator_class"])
print("Response: ", response["replies"][0].text)
Output:
WARNING:haystack.components.generators.chat.fallback:ChatGenerator AnthropicChatGenerator failed with error: Request timed out or interrupted...
WARNING:haystack.components.generators.chat.fallback:ChatGenerator GoogleGenAIChatGenerator failed with error: Error in Google Gen AI chat generation: 404 NOT_FOUND...
Successful ChatGenerator: OpenAIChatGenerator
Response: In "The Shawshank Redemption," ....
Tool and Toolset in AgentsYou can now combine both Tool and Toolset objects in the same tools list for Agent and ToolInvoker components. This update brings more flexibility, letting you organize tools into logical groups while still adding standalone tools in one go.
from haystack.components.agents import Agent
from haystack.tools import Tool, Toolset
math_toolset = Toolset([add_tool, multiply_tool])
weather_toolset = Toolset([weather_tool, forecast_tool])
agent = Agent(
chat_generator=generator,
tools=[math_toolset, weather_toolset, calendar_tool], # ✨ Now supported!
)
Tool and Toolset objects can now perform initialization during Agent or ToolInvoker warmup. This allows setup tasks such as connecting to databases, loading models, or initializing connection pools before the first use.
from haystack.tools import Toolset
from haystack.components.agents import Agent
# Custom toolset with initialization needs
class DatabaseToolset(Toolset):
def __init__(self, connection_string):
self.connection_string = connection_string
self.pool = None
super().__init__([query_tool, update_tool])
def warm_up(self):
# Initialize connection pool
self.pool = create_connection_pool(self.connection_string)
Updated our serialization and deserialization of PipelineSnapshots to work with python Enum classes.
Added FallbackChatGenerator that automatically retries different chat generators and returns first successful response with detailed information about which providers were tried.
Added pipeline_snapshot and pipeline_snapshot_file_path parameters to BreakpointException to provide more context when a pipeline breakpoint is triggered.
Added pipeline_snapshot_file_path parameter to PipelineRuntimeError to include a reference to the stored pipeline snapshot so it can be easily found.
A new component RegexTextExtractor which allows to extract text from chat messages or strings input based on custom regex pattern.
CSVToDocument: add conversion_mode='row' with optional content_column; each row becomes a Document; remaining columns stored in meta; default 'file' mode preserved.
Added the ability to resume an Agent from an AgentSnapshot while specifying a new breakpoint in the same run call. This allows stepwise debugging and precise control over chat generator inputs tool inputs before execution, improving flexibility when inspecting intermediate states. This addresses a previous limitation where passing both a snapshot and a breakpoint simultaneously would throw an exception.
Introduce SentenceTransformersSparseTextEmbedder and SentenceTransformersSparseDocumentEmbedder components. These components embed text and documents using sparse embedding models compatible with Sentence Transformers. Sparse embeddings are interpretable, efficient when used with inverted indexes, combine classic information retrieval with neural models, and are complementary to dense embeddings. Currently, the produced SparseEmbedding objects are compatible with the QdrantDocumentStore.
Usage example:
from haystack.components.embedders import SentenceTransformersSparseTextEmbedder
text_embedder = SentenceTransformersSparseTextEmbedder()
text_embedder.warm_up()
print(text_embedder.run("I love pizza!"))
# {'sparse_embedding': SparseEmbedding(indices=[999, 1045, ...], values=[0.918, 0.867, ...])}
Added a warm_up() function to the Tool dataclass, allowing tools to perform resource-intensive initialization before execution. Tools and Toolsets can now override the warm_up() method to establish connections to remote services, load models, or perform other preparatory operations. The ToolInvoker and Agent automatically call warm_up() on their tools during their own warm-up phase, ensuring tools are ready before use.
Fixed a serialization issue related to function objects in a pipeline; now they are converted to type None (functions cannot be serialized). This was preventing the successful setting of breakpoints in agents and their use as a resume point. If an error occurs during an Agent execution, for instance, during tool calling. In that case, a snapshot of the last successful step is raised, allowing the caller to catch it to inspect the possible reason for the crash and use it to resume the pipeline execution from that point onwards.
tools to agent run parameters to enhance the agent's flexibility. Users can now choose a subset of tools for the agent at runtime by providing a list of tool names, or supply an entirely new set by passing Tool objects or a Toolset.tools parameter across all tool-accepting components (Agent, ToolInvoker, OpenAIChatGenerator, AzureOpenAIChatGenerator, HuggingFaceAPIChatGenerator, HuggingFaceLocalChatGenerator) to accept either a mixed list of Tool and Toolset objects or just a Toolset object. Previously, components required either a list of Tool objects OR a single Toolset, but not both in the same list. Now users can organize tools into logical Toolsets while also including standalone Tool objects, providing greater flexibility in tool organization. For example: Agent(chat_generator=generator, tools=[math_toolset, weather_toolset, standalone_tool]). This change is fully backward compatible and preserves structure during serialization/deserialization, enabling proper round-trip support for mixed tool configurations._save_pipeline_snapshot to consolidate try-except logic and added a raise_on_failure option to control whether save failures raise an exception or are logged. _create_pipeline_snapshot now wraps _serialize_value_with_schema in try-except blocks to prevent failures from non-serializable pipeline inputs.run_async method to correctly handle async streaming callbacks. This previously triggered errors due to a bug.AgentSnapshot.response_format to None in OpenAIChatGenerator by default which doesn't follow the API spec. We now omit the variable if response_format is not passed by the user.OpenAIChatGenerator is properly serialized when response_format in generation_kwargs is provided as a dictionary (for example, {"type": "json_object"}). Previously, this caused serialization errors.ComponentTool when using inputs_from_state. Previously, parameters were only removed from the schema if the state key and parameter name matched exactly. For example, inputs_from_state={"text": "text"} removed text as expected, but inputs_from_state={"state_text": "text"} did not. This is now resolved, and such cases work as intended.SentenceTransformersEmbeddingBackend to ensure unique embedding IDs by incorporating all relevant arguments.BreakpointException when a ToolBreakpoint with a specific tool_name is provided in an assistant chat message containing multiple tool calls.OpenAIChatGenerator implementation uses ChatCompletionMessageCustomToolCall, which is only available in OpenAI client >=1.99.2. We now require openai>=1.99.2.@anakin87, @bilgeyucel, @davidsbatista, @dfokina, @Ryzhtus, @sjrl, @srini047, @tstadel, @vblagoje, @xoaryaa
Nothing published for this version
Added tools to agent run parameters to enhance the agent's flexibility. Users can now choose a subset of tools for the agent at runtime by providing a
run_async method to correctly handle async streaming callbacks. This previously triggered errors due to a bug.AgentSnapshot.response_format to None in OpenAIChatGenerator by default which doesn't follow the API spec. We now omit the variable if response_format is not passed by the user.Pipelines now capture a snapshot of the last successful step when a run fails, including intermediate outputs. This lets you diagnose issues (e.g., fa
Pipelines now capture a snapshot of the last successful step when a run fails, including intermediate outputs. This lets you diagnose issues (e.g., failed tool calls), fix them, and resume from the checkpoint instead of restarting the entire run. Currently supported for synchronous Pipeline and Agent (not yet in AsyncPipeline)
The snapshot is part of the exception raised with the PipelineRuntimeError when the pipeline run fails. You need to wrap your pipeline.run() in a try-except block.
try:
pipeline.run(data=input_data)
except PipelineRuntimeError as exc_info
snapshot = exc_info.value.pipeline_snapshot
intermediate_outputs = pipeline_snapshot.pipeline_state.pipeline_outputs
# Snapshot can be used to resume the execution of a Pipeline by passing it to the run() method using the snapshot argument
pipeline.run(data={}, snapshot=saved_snapshot)
OpenAIChatGenerator and AzureOpenAIChatGenerator support structured outputs via response_format (Pydantic model or JSON schema).
from pydantic import BaseModel
from haystack.components.generators.chat.openai import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
class CalendarEvent(BaseModel):
event_name: str
event_date: str
event_location: str
generator = OpenAIChatGenerator(generation_kwargs={"response_format": CalendarEvent})
message = "The Open NLP Meetup is going to be in Berlin at deepset HQ on September 19, 2025"
result = generator.run([ChatMessage.from_user(message)])
print(result["replies"][0].text)
# {"event_name":"Open NLP Meetup","event_date":"September 19","event_location":"deepset HQ, Berlin"}
PipelineToolThe new PipelineTool lets you expose entire Haystack Pipelines as LLM-compatible tools. It simplifies the previous SuperComponent + ComponentTool pattern into a single abstraction and directly exposes input_mapping and output_mapping for fine-grained control.
from haystack import Pipeline
from haystack.tools import PipelineTool
retrieval_pipeline = Pipeline()
retrieval_pipeline.add_component...
..
retrieval_tool = PipelineTool(
pipeline=retrieval_pipeline,
input_mapping={"query": ["bm25_retriever.query"]},
output_mapping={"ranker.documents": "documents"},
name="retrieval_tool",
description="Use to retrieve documents",
)
Agent’s system_prompt can now be updated dynamically at runtime for more flexible behavior.
OpenAIChatGenerator and AzureOpenAIChatGenerator now support structured outputs using response_format parameter that can be passed in generation_kwargs. The response_format parameter can be a Pydantic model or a JSON schema for non-streaming responses. For streaming responses, the response_format must be a JSON schema. Example usage of the response_format parameter:
from pydantic import BaseModel
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
class NobelPrizeInfo(BaseModel):
recipient_name: str
award_year: int
category: str
achievement_description: str
nationality: str
client = OpenAIChatGenerator(
model="gpt-4o-2024-08-06",
generation_kwargs={"response_format": NobelPrizeInfo}
)
response = client.run(messages=[
ChatMessage.from_user("In 2021, American scientist David Julius received the Nobel Prize in"
" Physiology or Medicine for his groundbreaking discoveries on how the human body"
" senses temperature and touch.")
])
print(response["replies"][0].text)
>>> {"recipient_name":"David Julius","award_year":2021,"category":"Physiology or Medicine","achievement_description":"David Julius was awarded for his transformative findings regarding the molecular mechanisms underlying the human body's sense of temperature and touch. Through innovative experiments, he identified specific receptors responsible for detecting heat and mechanical stimuli, ranging from gentle touch to pain-inducing pressure.","nationality":"American"}
Added PipelineTool, a new tool wrapper that allows Haystack Pipelines to be exposed as LLM-compatible tools.
SuperComponent and then passing it to ComponentTool.PipelineTool streamlines that pattern into a dedicated abstraction. It uses the same approach under the hood but directly exposes input_mapping and output_mapping so users can easily control which pipeline inputs and outputs are made available.Agent, enabling seamless integration of full pipelines as tools in multi-step reasoning workflows.Add a reasoning field to StreamingChunk that optionally takes in a ReasoningContent dataclass. This is to allow a structured way to pass reasoning contents to streaming chunks.
If an error occurs during the execution of a pipeline, the pipeline will raise an PipelineRuntimeError exception containing an error message and the components outputs up to the point of failure. This allows you to inspect and debug the pipeline up to the point of failure.
LinkContentFetcher: add request_headers to allow custom per-request HTTP headers. Header precedence: httpx client defaults → component defaults → request_headers → rotating User-Agent. Also make HTTP/2 handling import-safe: if h2 isn’t installed, fall back to HTTP/1.1 with a warning. Thanks @xoaryaa. (Fixes #9064)
A snapshot of the last successful step is also raised when an error occurs during a Pipeline run. Allowing the caller to catch it to inspect the possible reason for crash and use it to resume the pipeline execution from that point onwards.
Add exclude_subdomains parameter to SerperDevWebSearch component. When set to True, this parameter restricts search results to only the exact domains specified in allowed_domains, excluding any subdomains. For example, with allowed_domains=\["example.com"\] and exclude_subdomains=True, results from "blog.example.com" or "shop.example.com" will be filtered out, returning only results from "example.com". The parameter defaults to False to maintain backward compatibility with existing behavior.
system_prompt to agent run parameters to enhance customization and control over agent behavior.ChatMessage with invalid content parts. While LLMs may still generate messages in the wrong format, this error provides guidance on the expected structure, making retries easier and more reliable during agent runs. The error message was unintentionally removed during a previous refactoring.SentenceSplitter are now included in the distribution. They were previously missing due to a config in the .gitignore file.lambda_threshold=0.0 in SentenceTransformersDiversityRanker instead of overriding it with 0.5 due to short-circuit evaluation.MetaFieldGroupingRanker to still work when subgroup_by values are unhashable types like list. We handle this by stringfying the contents of doc.meta\[subgroup_by\] in the same we do this for values of doc.meta\[group_by\].ToolInvoker.run() to propagate contextvars into ThreadPoolExecutor workers, ensuring all tool spans (ComponentTool, Agent wrapped in ComponentTool, or custom tools) are correctly linked to the outer Agent's trace instead of starting new root traces. This improves end-to-end observability across the entire tool execution chain.from_dict method of MetadataRouter so the output_type parameter introduced in Haystack 2.17 is now optional when loading from YAML. This ensures compatibility with older Haystack pipelines.OpenAIChatGenerator, improved the logic to exclude unsupported custom tool calls. The previous implementation caused compatibility issues with the Mistral Haystack core integration, which extends OpenAIChatGenerator.ComponentTool when using inputs_from_state. Previously, parameters were only removed from the schema if the state key and parameter name matched exactly. For example, inputs_from_state={"text": "text"} removed text as expected, but inputs_from_state={"state_text": "text"} did not. This is now resolved, and such cases work as intended.@Amnah199, @Ujjwal-Bajpayee, @abdokaseb, @anakin87, @davidsbatista, @dfokina, @rigved-telang, @sjrl, @tstadel, @vblagoje, @xoaryaa
Nothing published for this version
Fixed the from_dict method of MetadataRouter so the output_type parameter introduced in Haystack 2.17 is now optional when loading from YAML. This ens
from_dict method of MetadataRouter so the output_type parameter introduced in Haystack 2.17 is now optional when loading from YAML. This ensures compatibility with older Haystack pipelines.OpenAIChatGenerator, improved the logic to exclude unsupported custom tool calls. The previous implementation caused compatibility issues with the Mistral Haystack core integration, which extends OpenAIChatGenerator.Following the introduction of image support in Haystack 2.16.0, we've expanded this to more model providers in Haystack and Haystack Core integrations
Following the introduction of image support in Haystack 2.16.0, we've expanded this to more model providers in Haystack and Haystack Core integrations.
Now supported: Amazon Bedrock, Anthropic, Azure, Google, Hugging Face API, Meta Llama API, Mistral, Nvidia, Ollama, OpenAI, OpenRouter, STACKIT.
We've improved several components to make them more flexible:
MetadataRouter, which is used to route Documents based on metadata, has been extended to also support routing ByteStream objects.SentenceWindowRetriever, which retrieves neighboring sentences around relevant Documents to provide full context, is now more flexible. Previously, its source_id_meta_field parameter accepted only a single field containing the ID of the original document. It now also accepts a list of fields, so that only documents matching all of the specified meta fields will be retrieved.MultiFileConverter outputs a new key failed in the result dictionary, which contains a list of files that failed to convert. The documents output is included only if at least one file is successfully converted. Previously, documents could still be present but empty if a file with a supported MIME type was provided but did not actually exist.
The finish_reason field behavior in HuggingFaceAPIChatGenerator has been updated. Previously, the new finish_reason mapping (introduced in Haystack 2.15.0 release) was only applied when streaming was enabled. When streaming was disabled, the old finish_reason was still returned. This change ensures the updated finish_reason values are consistently returned regardless of streaming mode.
How to know if you're affected: If you rely on finish_reason in responses from HuggingFaceAPIChatGenerator with streaming disabled, you may see different values after this upgrade.
What to do: Review the updated mapping:
length → lengtheos_token → stopstop_sequence → stoptool_callslist[Documents] or list[ByteStream] based on metadata.| (added in python 3.10) in serialize_type and Pipeline.connect(). These functions support both the typing.Union and | operators and mixtures of them for backwards compatibility.ReasoningContent as a new content part to the ChatMessage dataclass. This allows storing model reasoning text and additional metadata in assistant messages. Assistant messages can now include reasoning content using the reasoning parameter in ChatMessage.from_assistant(). We will progressively update the implementations for Chat Generators with LLMs that support reasoning to use this new content part.SentenceWindowRetriever's source_id_meta_field parameter to also accept a list of strings. If a list of fields are provided, then only documents matching both fields will be retrieved.HuggingFaceAPIChatGenerator to enable vision-language model (VLM) usage with images and text. Users can now send both text and images to VLM models through Hugging Face APIs. The implementation follows the HF VLM API format specification and maintains full backward compatibility with text-only messages.TextContent and ImageContent parts of ChatMessage.typing module.ChatMessage in Agent state schema validation. The validation now checks for issubclass(args[0], ChatMessage) instead of requiring exact type equality, allowing custom ChatMessage subclasses to be used in the messages field.ToolInvoker run method now accepts a list of tools. When provided, this list overrides the tools set in the constructor, allowing you to switch tools at runtime in previously built pipelines.The English and German abbreviation files used by the SentenceSplitter are now included in the distribution. They were previously missing due to a config in the .gitignore file.
Add encoding format keyword argument to OpenAI client when creating embeddings.
Addressed incorrect assumptions in the ChatMessage class that raised errors in valid usage scenario.
ChatMessage.from_user with content_parts: Previously, at least one text part was required, even though some model providers support messages with only image parts. This restriction has been removed. If a provider has such a limitation, it should now be enforced in the provider's implementation.
ChatMessage.to_openai_dict_format: Messages containing multiple text parts weren't supported, despite this being allowed by the OpenAI API. This has now been corrected.
Improved validation in the ChatMessage.from_user class method. The method now raises an error if neither text nor content_parts are provided. It does not raise an error if text is an empty string.
Ensure that the score field in SentenceTransformersSimilarityRanker is returned as a Python float instead of numpy.float32. This prevents potential serialization issues in downstream integrations.
Raise a RuntimeError when AsyncPipeline.run is called from within an async context, indicating that run_async should be used instead.
Prevented in-place mutation of input Document objects in all Extractor and Classifier components by creating copies with dataclasses.replace before processing.
Prevented in-place mutation of input Document objects in all DocumentEmbedder components by creating copies with dataclasses.replace before processing.
FileTypeRouter has a new parameter raise_on_failure with default value to False. When set to True, FileNotFoundError is always raised for non-existent files. Previously, this exception was raised only when processing a non-existent file and the meta parameter was provided to run().
Return a more informative error message when attempting to connect two components and the sender component does not have any OutputSockets defined.
Fix tracing context not propagated to tools when running via ToolInvoker.run_async
Ensure consistent behavior in SentenceTransformersDiversityRanker. Like other rankers, it now returns all documents instead of raising an error when top_k exceeds the number of available documents.
@abdokaseb @Amnah199 @anakin87 @bilgeyucel @ChinmayBansal @datbth @davidsbatista @dfokina @LastRemote @mpangrazzi @RafaelJohn9 @rolshoven @SaraCalla @SaurabhLingam @sjrl
Nothing published for this version
Nothing published for this version
Improved validation in the ChatMessage.from_user class method. The method now raises an error if neither text nor content_parts are provided. It does
ChatMessage.from_user class method. The method now raises an error if neither text nor content_parts are provided. It does not raise an error if text is an empty string.Nothing published for this version
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. T…
This release introduces Agent Breakpoints, a powerful new feature that enhances debugging and observability when working with Haystack Agents. You can pause execution mid-run by inserting breakpoints in the Agent or its tools to inspect internal state and resume execution seamlessly. This brings fine-grained control to agent development and significantly improves traceability during complex interactions.
from haystack.dataclasses.breakpoints import AgentBreakpoint, Breakpoint
from haystack.dataclasses import ChatMessage
chat_generator_breakpoint = Breakpoint(
component_name="chat_generator",
visit_count=0,
snapshot_file_path="debug_snapshots"
)
agent_breakpoint = AgentBreakpoint(break_point=chat_generator_breakpoint, agent_name='calculator_agent')
response = agent.run(
messages=[ChatMessage.from_user("What is 7 * (4 + 2)?")],
break_point=agent_breakpoint
)
You can now blend text and image capabilities across generation, indexing, and retrieval in Haystack.
New ImageContent Dataclass: A dedicated structure to store image data along with base64_image, mime_type, detail, and metadata.
Image-Aware Chat Generators: Image inputs are now supported in OpenAIChatGenerator
from haystack.dataclasses import ImageContent, ChatMessage
from haystack.components.generators.chat import OpenAIChatGenerator
image_url = "https://cdn.britannica.com/79/191679-050-C7114D2B/Adult-capybara.jpg"
image_content = ImageContent.from_url(image_url)
message = ChatMessage.from_user(
content_parts=["Describe the image in short.", image_content]
)
llm = OpenAIChatGenerator(model="gpt-4o-mini")
print(llm.run([message])["replies"][0].text)
Powerful Multimodal Components:
PDFToImageContent, ImageFileToImageContent, DocumentToImageContent: Convert PDFs, image files, and Documents into ImageContent objects.LLMDocumentContentExtractor: Extract text from images using a vision-enabled LLM.SentenceTransformersDocumentImageEmbedder: Generate embeddings from image-based documents using models like CLIP.DocumentLengthRouter: Route documents based on textual content length—ideal for distinguishing scanned PDFs from text-based ones.DocumentTypeRouter: Route documents automatically based on MIME type metadata.Prompt Building with Image Support: The ChatPromptBuilder now supports templates with embedded images, enabling dynamic multimodal prompt creation.
With these additions, you can now build multimodal agents and RAG pipelines that reason over both text and visual content, unlocking richer interactions and retrieval capabilities.
👉 Learn more about multimodality in our Introduction to Multimodal Text Generation.
Add to_dict and from_dict to ByteStream so it is consistent with our other dataclasses in having serialization and deserialization methods.
Add to_dict and from_dict to classes StreamingChunk, ToolCallResult, ToolCall, ComponentInfo, and ToolCallDelta to make it consistent with our other dataclasses in having serialization and deserialization methods.
Added the tool_invoker_kwargs param to Agent so additional kwargs can be passed to the ToolInvoker like max_workers and enable_streaming_callback_passthrough.
ChatPromptBuilder now supports special string templates in addition to a list of ChatMessage objects. This new format is more flexible and allows structured parts like images to be included in the templatized ChatMessage.
from haystack.components.builders import ChatPromptBuilder
from haystack.dataclasses.chat_message import ImageContent
template = """
{% message role="user" %}
Hello! I am {{user_name}}.
What's the difference between the following images?
{% for image in images %}
{{ image | templatize_part }}
{% endfor %}
{% endmessage %}
"""
images=[
ImageContent.from_file_path("apple-fruit.jpg"),
ImageContent.from_file_path("apple-logo.jpg")
]
builder = ChatPromptBuilder(template=template) builder.run(user_name="John", images=images)
Added convenience class methods to the ImageContent dataclass to create ImageContent objects from file paths and URLs.
Added multiple converters to help convert image data between different formats:
DocumentToImageContent: Converts documents sourced from PDF and image files into ImageContents.
ImageFileToImageContent: Converts image files to ImageContent objects.
ImageFileToDocument: Converts image file references into empty Document objects with associated metadata.
PDFToImageContent: Converts PDF files to ImageContent objects.
Chat Messages with the user role can now include images using the new ImageContent dataclass. We've added image support to OpenAIChatGenerator, and plan to support more model providers over time.
Raise a warning when a pipeline can no longer proceed because all remaining components are blocked from running and no expected pipeline outputs have been produced. This scenario can occur legitimately. For example, in pipelines with mutually exclusive branches where some components are intentionally blocked. To help avoid false positives, the check ensures that none of the expected outputs (as defined by Pipeline().outputs()) have been generated during the current run.
Added source_id_meta_field and split_id_meta_field to SentenceWindowRetriever for customizable metadata field names. Added raise_on_missing_meta_fields to control whether a ValueError is raised if any of the documents at runtime are missing the required meta fields (set to True by default). If False, then the documents missing the meta field will be skipped when retrieving their windows, but the original document will still be included in the results.
Add a ComponentInfo dataclass to the haystack.dataclasses module. This dataclass is used to store information about the component. We pass it to StreamingChunk so we can tell from which component a stream is coming from.
Pass the component_info to the StreamingChunk in the OpenAIChatGenerator, AzureOpenAIChatGenerator, HuggingFaceAPIChatGenerator and HuggingFaceLocalChatGenerator.
Added the enable_streaming_callback_passthrough to the ToolInvoker init, run and run_async methods. If set to True the ToolInvoker will try and pass the streaming_callback function to a tool's invoke method only if the tool's invoke method has streaming_callback in its signature.
Added new HuggingFaceTEIRanker component to enable reranking with Text Embeddings Inference (TEI) API. This component supports both self-hosted Text Embeddings Inference services and Hugging Face Inference Endpoints.
Added a raise_on_failure boolean parameter to OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder. If set to True then the component will raise an exception when there is an error with the API request. It is set to False by default to so the previous behavior of logging an exception and continuing is still the default.
ToolInvoker now executes tool_calls in parallel for both sync and async mode.
Add AsyncHFTokenStreamingHandler for async streaming support in HuggingFaceLocalChatGenerator
Updated StreamingChunk to add the fields tool_calls, tool_call_result, index, and start to make it easier to format the stream in a streaming callback.
HuggingFaceAPIGenerator might no longer work with the Hugging Face Inference API. As of July 2025, the Hugging Face Inference API no longer offers generative models that support the text_generation endpoint. Generative models are now only available through providers that support the chat_completion endpoint. As a result, the HuggingFaceAPIGenerator component might not work with the Hugging Face Inference API. It still works with Hugging Face Inference Endpoints and self-hosted TGI instances. To use generative models via Hugging Face Inference API, please use the HuggingFaceAPIChatGenerator component, which supports the chat_completion endpoint.
All parameters of the Pipeline.draw() and Pipeline.show() methods must now be specified as keyword arguments. Example:
pipeline.draw(
path="output.png",
server_url="https://custom-server.com",
params=None,
timeout=30,
super_component_expansion=False
)
The deprecated async_executor parameter has been removed from the ToolInvoker class. Please use the max_workers parameter instead and a ThreadPoolExecutor with these workers will be created automatically for parallel tool invocations.
The deprecated State class has been removed from the haystack.dataclasses module. The State class is now part of the haystack.components.agents module.
Remove the deserialize_value_with_schema_legacy function from the base_serialization module. This function was used to deserialize State objects created with Haystack 2.14.0 or older. Support for the old serialization format is removed in Haystack 2.16.0.
Add guess_mime_type parameter to Bytestream.from_file_path()
Add the init parameter skip_empty_documents to the DocumentSplitter component. The default value is True. Setting it to False can be useful when downstream components in the Pipeline (like LLMDocumentContentExtractor) can extract text from non-textual documents.
Test that our type validation and connection validation works with builtin python types introduced in 3.9. We found that these types were already supported, we just now add explicit tests for them.
We relaxed the requirement that in ToolCallDelta (introduced in Haystack 2.15) which required the parameters arguments or name to be populated to be able to create a ToolCallDelta dataclass. We remove this requirement to be more in line with OpenAI's SDK and since this was causing errors for some hosted versions of open source models following OpenAI's SDK specification.
Added return_embedding parameter inside InMemoryDocumentStore::init method.
Updated methods bm25_retrieval, and filter_documents to use self.return_embedding to determine whether embeddings are returned.
Updated tests (test_in_memory & test_in_memory_embedding_retriever) to reflect the changes in the InMemoryDocumentStore.
Made doc-parser a core dependency since ComponentTool that uses it is one of the core Tool components.
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. This method is useful for checking that all required connections in a pipeline have a connection and is automatically called in the run method of Pipeline. It is being exposed as public for users who would like to call this method before runtime to validate the pipeline.
Haystack's core modules are now ["type complete"](https://typing.python.org/en/latest/guides/libraries.html#how-much-of-my-library-needs-types), meaning that all function parameters and return types are explicitly annotated. This increases the usefulness of the newly added py.typed marker and sidesteps differences in type inference between the various type checker implementations.
Refactors the HuggingFaceAPIChatGenerator to use the util method _convert_streaming_chunks_to_chat_message. This is to help with being consistent for how we convert StreamingChunks into a final ChatMessage.
We also add ComponentInfo to the StreamingChunks made in HuggingFaceGenerator, and HugginFaceLocalGenerator so we can tell from which component a stream is coming from.
If only system messages are provided as input a warning will be logged to the user indicating that this likely not intended and that they should probably also provide user messages.
Fix _convert_streaming_chunks_to_chat_message which is used to convert Haystack StreamingChunks into a Haystack ChatMessage. This fixes the scenario where one StreamingChunk contains two ToolCallDeltas in StreamingChunk.tool_calls. With this fix this correctly saves both ToolCallDeltas whereas before they were overwriting each other. This only occurs with some LLM providers like Mistral (and not OpenAI) due to how the provider returns tool calls.
Fixed a bug in the print_streaming_chunk utility function that prevented tool call name from being printed.
Fix component_invoker used by ComponentTool to work when a dataclass like ChatMessage is directly passed to component_tool.invoke(...). Previously this would either cause an error or silently skip your input.
RecursiveDocumentSplitter now generates a unique Document.id for every chunk. The meta fields (split_id, parent_id, etc.) are populated [before]() Document creation, so the hash used for id generation is always unique.
In ConditionalRouter fixed the to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. This occurs when a user specifies a route in ConditionalRouter to have multiple outputs.
When calling set_output_types we now also check that the decorator @component.output_types is not present on the run_async method of a Component. Previously we only checked that the Component.run method did not possess the decorator.
@Amnah199 @RafaelJohn9 @anakin87 @bilgeyucel @davidsbatista @julian-risch @kanenorman @kr1shnasomani @mathislucka @mpangrazzi @sjrl @srishti-git1110
Nothing published for this version
We’ve relaxed the requirements for the ToolCallDelta dataclass (introduced in Haystack 2.15). Previously, creating a ToolCallDelta instance required e
ToolCallDelta dataclass (introduced in Haystack 2.15). Previously, creating a ToolCallDelta instance required either the parameters argument or the name to be set. This constraint has now been removed to align more closely with OpenAI's SDK behavior.
The change was necessary as the stricter requirement was causing errors in certain hosted versions of open-source models that adhere to the OpenAI SDK specification.print_streaming_chunk utility function that prevented ToolCall name from being printed.Nothing published for this version
Fix _convert_streaming_chunks_to_chat_message which is used to convert Haystack StreamingChunks into a Haystack ChatMessage. This fixes the scenario w
_convert_streaming_chunks_to_chat_message which is used to convert Haystack StreamingChunks into a Haystack ChatMessage. This fixes the scenario where one StreamingChunk contains two ToolCallDetlas in StreamingChunk.tool_calls. With this fix this correctly saves both ToolCallDeltas whereas before they were overwriting each other. This only occurs with some LLM providers like Mistral (and not OpenAI) due to how the provider returns tool calls.Nothing published for this version
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. T…
ToolInvoker now processes all tool calls passed to run or run_async in parallel using an internal ThreadPoolExecutor. This improves performance by reducing the time spent on sequential tool invocations.ToolInvoker to batch and process multiple tool calls concurrently, allowing Agents to run complex pipelines efficiently with decreased latency.async_executor. ToolInvoker manages its own executor, configurable via the max_workers parameter in init.The new LLMMessagesRouter component that classifies and routes incoming ChatMessage objects to different connections using a generative LLM. This component can be used with general-purpose LLMs and with specialized LLMs for moderation like Llama Guard.
Usage example:
from haystack.components.generators.chat import HuggingFaceAPIChatGenerator
from haystack.components.routers.llm_messages_router import LLMMessagesRouter
from haystack.dataclasses import ChatMessage
chat_generator = HuggingFaceAPIChatGenerator(api_type="serverless_inference_api", api_params={"model": "meta-llama/Llama-Guard-4-12B", "provider": "groq"}, )
router = LLMMessagesRouter(chat_generator=chat_generator, output_names=["unsafe", "safe"], output_patterns=["unsafe", "safe"])
print(router.run([ChatMessage.from_user("How to rob a bank?")]))
HuggingFaceTEIRanker enables end-to-end reranking via the Text Embeddings Inference (TEI) API. It supports both self-hosted TEI services and Hugging Face Inference Endpoints, giving you flexible, high-quality reranking out of the box.
Added a ComponentInfo dataclass to haystack to store information about the component. We pass it to StreamingChunk so we can tell from which component a stream is coming.
Pass the component_info to the StreamingChunk in the OpenAIChatGenerator, AzureOpenAIChatGenerator, HuggingFaceAPIChatGenerator, HuggingFaceGenerator, HugginFaceLocalGenerator and HuggingFaceLocalChatGenerator.
Added the enable_streaming_callback_passthrough to the init, run and run_async methods of ToolInvoker. If set to True the ToolInvoker will try and pass the streaming_callback function to a tool's invoke method only if the tool's invoke method has streaming_callback in its signature.
Added dedicated finish_reason field to StreamingChunk class to improve type safety and enable sophisticated streaming UI logic. The field uses a FinishReason type alias with standard values: "stop", "length", "tool_calls", "content_filter", plus Haystack-specific value "tool_call_results" (used by ToolInvoker to indicate tool execution completion).
Updated ToolInvoker component to use the new finish_reason field when streaming tool results. The component now sets finish_reason="tool_call_results" in the final streaming chunk to indicate that tool execution has completed, while maintaining backward compatibility by also setting the value in meta["finish_reason"].
Added a raise_on_failure boolean parameter to OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder. If set to True then the component will raise an exception when there is an error with the API request. It is set to False by default so the previous behavior of logging an exception and continuing is still the default.
Add AsyncHFTokenStreamingHandler for async streaming support in HuggingFaceLocalChatGenerator
For HuggingFaceAPIGenerator and HuggingFaceAPIChatGenerator all additional key, value pairs passed in api_params are now passed to the initializations of the underlying Inference Clients. This allows passing of additional parameters to the clients like timeout, headers, provider, etc. This means we now can easily specify a different inference provider by passing the provider key in api_params.
Updated StreamingChunk to add the fields tool_calls, tool_call_result, index, and start to make it easier to format the stream in a streaming callback.
ToolCallDelta for the StreamingChunk.tool_calls field to reflect that the arguments can be a string delta.print_streaming_chunk and _convert_streaming_chunks_to_chat_message utility methods to use these new fields. This especially improves the formatting when using print_streaming_chunk with Agent.OpenAIGenerator, OpenAIChatGenerator, HuggingFaceAPIGenerator, HuggingFaceAPIChatGenerator, HuggingFaceLocalGenerator and HuggingFaceLocalChatGenerator to follow the new dataclasses.ToolInvoker to follow the StreamingChunk dataclass.Added a new deserialize_component_inplace function to handle generic component deserialization that works with any component type.
Made doc-parser a core dependency since ComponentTool that uses it is one of the core Tool components.
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. This method is useful for checking that all required connections in a pipeline have a connection and is automatically called in the run method of Pipeline. It is being exposed as public for users who would like to call this method before runtime to validate the pipeline.
For component run Datadog tracing, set the span resource name to the component name instead of the operation name.
Added a trust_remote_code parameter to the SentenceTransformersSimilarityRanker component. When set to True, this enables execution of custom models and scripts hosted on the Hugging Face Hub.
Add a new parameter require_tool_call_ids to ChatMessage.to_openai_dict_format. The default is True, for compatibility with OpenAI's Chat API: if the id field is missing in a Tool Call, an error is raised. Using False is useful for shallow OpenAI-compatible APIs, where the id field is not required.
Haystack's core modules are now "type complete", meaning that all function parameters and return types are explicitly annotated. This increases the usefulness of the newly added py.typed marker and sidesteps differences in type inference between the various type checker implementations.
HuggingFaceAPIChatGenerator now uses the util method _convert_streaming_chunks_to_chat_message. This is to help with being consistent for how we convert StreamingChunks into a final ChatMessage.
async_executor parameter in ToolInvoker is deprecated in favor of max_workers parameter and will be removed in Haystack 2.16.0. You can use max_workers parameter to control the number of threads used for parallel tool calling.to_dict and from_dict of ToolInvoker to properly serialize the streaming_callback init parameter.raise_on_failure=False and an error occurs mid-batch that the following embeddings would be paired with the wrong documents.ComponentTool to work when a dataclass like ChatMessage is directly passed to component_tool.invoke(...). Previously this would either cause an error or silently skip your input.LLMMetadataExtractor that occurred when processing Document objects with None or empty string content. The component now gracefully handles these cases by marking such documents as failed and providing an appropriate error message in their metadata, without attempting an LLM call.Document.id for every chunk. The meta fields (split_id, parent_id, etc.) are populated before Document creation, so the hash used for id generation is always unique.ConditionalRouter fixed the to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. This occurs when a user specifies a route in ConditionalRouter to have multiple outputs.GeneratedAnswer when ChatMessage objects are nested in meta.ComponentTool and Tool when specifying outputs_to_string. Previously an error occurred on deserialization right after serializing if outputs_to_string is not None.set_output_types we now also check that the decorator @component.output_types is not present on the run_async method of a Component. Previously we only checked that the Component.run method did not possess the decorator.is not with != when checking the type List[ChatMessage]. This prevents false mismatches due to Python's is operator comparing object identity instead of equality.__init__.py files. This ensures that short imports like from haystack.components.builders import ChatPromptBuilder work equivalently to from haystack.components.builders.chat_prompt_builder import ChatPromptBuilder, without causing errors or warnings in mypy/Pylance.SuperComponent class can now correctly serialize and deserialize a SuperComponent based on an async pipeline. Previously, the SuperComponent class always assumed the underlying pipeline was synchronous.OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder where if an OpenAI API error occurred mid-batch then the following embeddings would be paired with the wrong documents.Nothing published for this version
In ConditionalRouter fixed the to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. Thi
to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. This occurs when a user specifies a route in ConditionalRouter to have multiple outputs.outputs_to_string. Previously an error occurred on deserialization right after serializing if outputs_to_string is not None.Nothing published for this version
Fixed a bug in OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder where if an OpenAI API error occurred mid-batch then the following embeddings wo
OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder where if an OpenAI API error occurred mid-batch then the following embeddings would be paired with the wrong documents.raise_on_failure boolean parameter to OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder. If set to True then the component will raise an exception when there is an error with the API request. It is set to False by default so the previous behavior of logging an exception and continuing is still the default.Nothing published for this version
Fixed a mypy issue in the OpenAIChatGenerator and its handling of stream responses. This issue only occurs with mypy \>=1.16.0.
Nothing published for this version
This component replaces the legacy TransformersSimilarityRanker component, which may be deprecated in a future release, with removal following a depre…
We've improved agent workflows with better message handling and streaming support. Agent component now returns a last_message output for quick access to the final message, and can use a streaming_callback to emit tool results in real time. You can use the updated print_streaming_chunk or write your own callback function to enable ToolCall details during streaming.
from haystack.components.websearch import SerperDevWebSearch
from haystack.components.agents import Agent
from haystack.components.generators.utils import print_streaming_chunk
from haystack.tools import tool, ComponentTool
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
web_search = ComponentTool(name="web_search", component=SerperDevWebSearch(top_k=5))
wiki_search = ComponentTool(name="wiki_search", component=SerperDevWebSearch(top_k=5, allowed_domains=["https://www.wikipedia.org/"]))
research_agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-4o-mini"),
system_prompt="""
You are a research agent that can find information on web or specifically on wikipedia.
Use wiki_search tool if you need facts and use web_search tool for latest news on topics.
Use one tool at a time, use the other tool if the retrieved information is not enough.
Summarize the retrieved information before returning response to the user.
""",
tools=[web_search, wiki_search],
streaming_callback=print_streaming_chunk
)
result = research_agent.run(messages=[ChatMessage.from_user("Can you tell me about Florence Nightingale's life?")])
Enabling streaming with print_streaming_chunk function looks like this:
[TOOL CALL]
Tool: wiki_search
Arguments: {"query":"Florence Nightingale"}
[TOOL RESULT]
{'documents': [{'title': 'List of schools in Nottinghamshire', 'link': 'https://www.wikipedia.org/wiki/List_of_schools_in_Nottinghamshire', 'position': 1, 'id': 'a6d0fe00f1e0cd06324f80fb926ba647878fb7bee8182de59a932500aeb54a5b', 'content': 'The Florence Nightingale Academy, Eastwood; The Flying High Academy, Mansfield; Forest Glade Primary School, Sutton-in-Ashfield; Forest Town Primary School ...', 'blob': None, 'score': None, 'embedding': None, 'sparse_embedding': None}], 'links': ['https://www.wikipedia.org/wiki/List_of_schools_in_Nottinghamshire']}
...
Print the last_message
print("Final Answer:", result["last_message"].text)
>>> Final Answer: Florence Nightingale (1820-1910) was a pioneering figure in nursing and is often hailed as the founder of modern nursing. She was born...
Additionally, AnswerBuilder stores all generated messages in all_messages meta field of GeneratedAnswer and supports a new last_message_only mode for lightweight flows where only the final message needs to be processed.
We extended pipeline.draw() and pipeline.show(), which save pipeline diagrams to images files or display them in Jupyter notebooks. You can now pass super_component_expansion=True to expand any SuperComponents and draw more detailed visualizations.
Here is an example with a pipeline containing MultiFileConverter and DocumentPreprocssor SuperComponents. After installing the dependencies that the MultiFileConverter needs for all supported file formats via pip install haystack-ai pypdf markdown-it-py mdit_plain trafilatura python-pptx python-docx jq openpyxl tabulate pandas, you can run:
from pathlib import Path
from haystack import Pipeline
from haystack.components.converters import MultiFileConverter
from haystack.components.preprocessors import DocumentPreprocessor
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore
document_store = InMemoryDocumentStore()
pipeline = Pipeline()
pipeline.add_component("converter", MultiFileConverter())
pipeline.add_component("preprocessor", DocumentPreprocessor())
pipeline.add_component("writer", DocumentWriter(document_store = document_store))
pipeline.connect("converter", "preprocessor")
pipeline.connect("preprocessor", "writer")
# expanded pipeline that shows all components
path = Path("expanded_pipeline.png")
pipeline.draw(path=path, super_component_expansion=True)
# original pipeline
path = Path("original_pipeline.png")
pipeline.draw(path=path)
We added a new SentenceTransformersSimilarityRanker component that uses the Sentence Transformers library to rank documents based on their semantic similarity to the query. This component replaces the legacy TransformersSimilarityRanker component, which may be deprecated in a future release, with removal following a deprecation period. The SentenceTransformersSimilarityRanker also allows choosing different inference backends: PyTorch, ONNX, and OpenVINO. For example, after installing sentence-transformers>=4.1.0, you can run:
from haystack.components.rankers import SentenceTransformersSimilarityRanker
from haystack.utils.device import ComponentDevice
onnx_ranker = SentenceTransformersSimilarityRanker(
model="sentence-transformers/all-MiniLM-L6-v2",
token=None,
device=ComponentDevice.from_str("cpu"),
backend="onnx",
)
onnx_ranker.warm_up()
docs = [Document(content="Berlin"), Document(content="Sarajevo")]
output = onnx_ranker.run(query="City in Germany", documents=docs)
ranked_docs = output["documents"]
py.typed file to Haystack to enable type information to be used by downstream projects, in line with PEP 561. This means Haystack's type hints will now be visible to type checkers in projects that depend on it. Haystack is primarily type checked using mypy (not pyright) and, despite our efforts, some type information can be incomplete or unreliable. If you use static type checking in your own project, you may notice some changes: previously, Haystack's types were effectively treated as Any, but now actual type information will be available and enforced. We'll continue improving typing with the next release.deserialize_tools_inplace utility function has been removed. Use deserialize_tools_or_toolset_inplace instead, importing it as follows: from haystack.tools import deserialize_tools_or_toolset_inplace.Added run_async method to ToolInvoker class to allow asynchronous tool invocations.
Agent can now stream tool result with run_async method as well.
Introduced serialize_value and deserialize_value utility methods for consistent value (de)serialization across modules.
Moved the State class to the agents.state module and added serialization and deserialization capabilities.
Add support for multiple outputs in ConditionalRouter
Implement JSON-safe serialization for OpenAI usage data by converting token counts and details (like CompletionTokensDetails and PromptTokensDetails) into plain dictionaries.
Added a new SentenceTransformersSimilarityRanker component that uses the Sentence Transformers library to rank documents based on their semantic similarity to the query. This component is a replacement for the legacy TransformersSimilarityRanker component, which may be deprecated in a future release, with removal following after a deprecation period. The SentenceTransformersSimilarityRanker also allows choosing different inference backends: PyTorch, ONNX, and OpenVINO. To use the SentenceTransformersSimilarityRanker, you need to install sentence-transformers>=4.1.0.
Add a streaming_callback parameter to ToolInvoker to enable streaming of tool results. Note that tool_result is emitted only after the tool execution completes and is not streamed incrementally.
Update print_streaming_chunk to print ToolCall information if it is present in the chunk's metadata.
Update Agent to forward the streaming_callback to ToolInvoker to emit tool results during tool invocation.
Enhance SuperComponent's type compatibility check to return the detected common type between two input types.
When using HuggingFaceAPIChatGenerator with streaming, the returned ChatMessage now contains the number of prompt tokens and completion tokens in its meta data. Internally, the HuggingFaceAPIChatGenerator requests an additional streaming chunk that contains usage data. It then processes the usage streaming chunk to add usage meta data to the returned ChatMessage.
We now have a Protocol for TextEmbedder. The protocol makes it easier to create custom components or SuperComponents that expect any TextEmbedder as init parameter.
We added a Component signature validation method that details the mismatches between the run and run_async method signatures. This allows a user to debug custom components easily.
Enhanced the AnswerBuilder component with two agent-friendly features:
meta field of the GeneratedAnswer objects under an all_messages key, improving traceability and debugging capabilities.last_message_only parameter that, when set to True, processes only the last message in the replies while still preserving the complete conversation history in metadata. This is particularly useful for agent workflows where only the final response needs to be processed.A variety of improvements have been made so an Agent component can be directly used in ComponentTool enabling straightforward building of Multi-Agent systems. These improvements include:
last_message field to the Agent's output which returns the last generated ChatMessage._default_output_handler in the ToolInvoker to try and first serialize the outputs in the tool result before converting it into a string. This is especially relevant for getting a better representation when stringifying dataclasses like ChatMessage.Added type hints to the component decorator. This improves support for Pyright/Pylance, enabling IDEs like VSCode to show docstrings for components.
Updated pipeline execution logic to use a new utility method _deepcopy_with_exceptions, which attempts to deep copy an object and safely falls back to the original object if copying fails. Additionally _deepcopy_with_exceptions skips deep-copying of Component, Tool, and Toolset instances when used as runtime parameters. This prevents errors and unintended behavior caused by trying to deepcopy objects that contain non-copyable attributes (e.g. Jinja2 templates, clients). Previously, standard deepcopy was used on inputs and outputs which occasionally lead to errors since certain Python objects cannot be deepcopied.
Refactored JSON Schema generation for ComponentTool parameters using Pydantic’s model_json_schema, enabling expanded type support (e.g. Union, Enum, Dict, etc.). We also convert dataclasses to Pydantic models before calling model_json_schema to preserve docstring descriptions of the parameters in the schema. This means dataclasses like ChatMessage, Document, etc. now have correctly defined JSON schemas.
The draw() and show() methods from Pipeline now have an extra boolean parameter, super_component_expansion, which, when set to True and if the pipeline contains SuperComponents, the visualisation diagram will show the internal structure of super-components as if they were components part of the pipeline instead of a "black-box" with the name of the SuperComponent.
Improve the type annotations for @component and the Component protocol. The type checker can now ensure that a @component class provides a compatible run() method, whose required return type has been changed from Dict[str, Any] (invariant) to the Mapping[str, Any] to allow TypedDict to be used for output types.
ComponentTool now preserves and combines docstrings from underlying pipeline components when wrapping a SuperComponent. When a SuperComponent is used with ComponentTool, two key improvements are made:
These changes make SuperComponents much more useful with LLM function calling as the LLM will get detailed information about both the component's purpose and its parameters.
Adds local_files_only parameter to SentenceTransformersDocumentEmbedder and SentenceTransformersTextEmbedder to allow loading models in offline mode.
The DocumentRecallEvaluator was updated. Now, when in MULTI_HIT mode, the division is over the unique ground truth documents instead of the total number of ground truth documents. We also added checks for emptiness. If there are no retrieved documents or all of them have an empty string as content, we return 0.0 and log a warning. Likewise, if there are no ground truth documents or all of them have an empty string as content, we return 0.0 and log a warning.
State class in the dataclasses module. Users are encouraged to transition to the new version of State now located in the agents.state module. A deprecation warning has been added to guide this migration.__deepcopy__ of ComponentTool to gracefully handle NotImplementedError when trying to deepcopy attributes.RecursiveDocumentSplitter was fixed for the case where a split_text is longer than the split_length and recursive chunking is triggered.huggingface_hub>=0.31.0. In the huggingface_hub library, arguments attribute of ChatCompletionInputFunctionDefinition has been renamed to parameters. Our implementation is compatible with both the legacy version and the new one.HuggingFaceAPIChatGenerator now checks the type of the arguments variable in the tool calls returned by the Hugging Face API. If arguments is a JSON string, it is parsed into a dictionary. Previously, the arguments type was not checked, which sometimes led to failures later in the tool workflow.ctx.run(...) so we can preserve context like the active tracing span. This now means if your component 1) only has a sync run method and 2) it logs something to the tracer then this trace will be properly nested within the parent context.LLMMetadataExtractor that occurred when processing Document objects with None or empty string content. The component now gracefully handles these cases by marking such documents as failed and providing an appropriate error message in their metadata, without attempting an LLM call.component_tool.invoke(...). Previously this would either cause an error or silently skip your input.Special thanks and congratulations to our first time contributors!
Full Changelog: https://github.com/deepset-ai/haystack/compare/v2.13.0...v2.14.0
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →