NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #5108 most downloaded on PyPI
LLM framework to build customizable, production-ready LLM applications. Connect components (models, vector DBs, file converters) to pipelines or agents that can interact with your data.
Last release today
05 Oct 2026
Ships fairly regularly
a new release about every 8 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
231 releases · first in 2023
Following the introduction of image support in Haystack 2.16.0, we've expanded this to more model providers in Haystack and Haystack Core integrations
Following the introduction of image support in Haystack 2.16.0, we've expanded this to more model providers in Haystack and Haystack Core integrations.
Now supported: Amazon Bedrock, Anthropic, Azure, Google, Hugging Face API, Meta Llama API, Mistral, Nvidia, Ollama, OpenAI, OpenRouter, STACKIT.
We've improved several components to make them more flexible:
MetadataRouter, which is used to route Documents based on metadata, has been extended to also support routing ByteStream objects.SentenceWindowRetriever, which retrieves neighboring sentences around relevant Documents to provide full context, is now more flexible. Previously, its source_id_meta_field parameter accepted only a single field containing the ID of the original document. It now also accepts a list of fields, so that only documents matching all of the specified meta fields will be retrieved.MultiFileConverter outputs a new key failed in the result dictionary, which contains a list of files that failed to convert. The documents output is included only if at least one file is successfully converted. Previously, documents could still be present but empty if a file with a supported MIME type was provided but did not actually exist.
The finish_reason field behavior in HuggingFaceAPIChatGenerator has been updated. Previously, the new finish_reason mapping (introduced in Haystack 2.15.0 release) was only applied when streaming was enabled. When streaming was disabled, the old finish_reason was still returned. This change ensures the updated finish_reason values are consistently returned regardless of streaming mode.
How to know if you're affected: If you rely on finish_reason in responses from HuggingFaceAPIChatGenerator with streaming disabled, you may see different values after this upgrade.
What to do: Review the updated mapping:
length → lengtheos_token → stopstop_sequence → stoptool_callslist[Documents] or list[ByteStream] based on metadata.| (added in python 3.10) in serialize_type and Pipeline.connect(). These functions support both the typing.Union and | operators and mixtures of them for backwards compatibility.ReasoningContent as a new content part to the ChatMessage dataclass. This allows storing model reasoning text and additional metadata in assistant messages. Assistant messages can now include reasoning content using the reasoning parameter in ChatMessage.from_assistant(). We will progressively update the implementations for Chat Generators with LLMs that support reasoning to use this new content part.SentenceWindowRetriever's source_id_meta_field parameter to also accept a list of strings. If a list of fields are provided, then only documents matching both fields will be retrieved.HuggingFaceAPIChatGenerator to enable vision-language model (VLM) usage with images and text. Users can now send both text and images to VLM models through Hugging Face APIs. The implementation follows the HF VLM API format specification and maintains full backward compatibility with text-only messages.TextContent and ImageContent parts of ChatMessage.typing module.ChatMessage in Agent state schema validation. The validation now checks for issubclass(args[0], ChatMessage) instead of requiring exact type equality, allowing custom ChatMessage subclasses to be used in the messages field.ToolInvoker run method now accepts a list of tools. When provided, this list overrides the tools set in the constructor, allowing you to switch tools at runtime in previously built pipelines.The English and German abbreviation files used by the SentenceSplitter are now included in the distribution. They were previously missing due to a config in the .gitignore file.
Add encoding format keyword argument to OpenAI client when creating embeddings.
Addressed incorrect assumptions in the ChatMessage class that raised errors in valid usage scenario.
ChatMessage.from_user with content_parts: Previously, at least one text part was required, even though some model providers support messages with only image parts. This restriction has been removed. If a provider has such a limitation, it should now be enforced in the provider's implementation.
ChatMessage.to_openai_dict_format: Messages containing multiple text parts weren't supported, despite this being allowed by the OpenAI API. This has now been corrected.
Improved validation in the ChatMessage.from_user class method. The method now raises an error if neither text nor content_parts are provided. It does not raise an error if text is an empty string.
Ensure that the score field in SentenceTransformersSimilarityRanker is returned as a Python float instead of numpy.float32. This prevents potential serialization issues in downstream integrations.
Raise a RuntimeError when AsyncPipeline.run is called from within an async context, indicating that run_async should be used instead.
Prevented in-place mutation of input Document objects in all Extractor and Classifier components by creating copies with dataclasses.replace before processing.
Prevented in-place mutation of input Document objects in all DocumentEmbedder components by creating copies with dataclasses.replace before processing.
FileTypeRouter has a new parameter raise_on_failure with default value to False. When set to True, FileNotFoundError is always raised for non-existent files. Previously, this exception was raised only when processing a non-existent file and the meta parameter was provided to run().
Return a more informative error message when attempting to connect two components and the sender component does not have any OutputSockets defined.
Fix tracing context not propagated to tools when running via ToolInvoker.run_async
Ensure consistent behavior in SentenceTransformersDiversityRanker. Like other rankers, it now returns all documents instead of raising an error when top_k exceeds the number of available documents.
@abdokaseb @Amnah199 @anakin87 @bilgeyucel @ChinmayBansal @datbth @davidsbatista @dfokina @LastRemote @mpangrazzi @RafaelJohn9 @rolshoven @SaraCalla @SaurabhLingam @sjrl
One column per month.
Nothing published for this version
Nothing published for this version
Improved validation in the ChatMessage.from_user class method. The method now raises an error if neither text nor content_parts are provided. It does
ChatMessage.from_user class method. The method now raises an error if neither text nor content_parts are provided. It does not raise an error if text is an empty string.Nothing published for this version
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. T…
This release introduces Agent Breakpoints, a powerful new feature that enhances debugging and observability when working with Haystack Agents. You can pause execution mid-run by inserting breakpoints in the Agent or its tools to inspect internal state and resume execution seamlessly. This brings fine-grained control to agent development and significantly improves traceability during complex interactions.
from haystack.dataclasses.breakpoints import AgentBreakpoint, Breakpoint
from haystack.dataclasses import ChatMessage
chat_generator_breakpoint = Breakpoint(
component_name="chat_generator",
visit_count=0,
snapshot_file_path="debug_snapshots"
)
agent_breakpoint = AgentBreakpoint(break_point=chat_generator_breakpoint, agent_name='calculator_agent')
response = agent.run(
messages=[ChatMessage.from_user("What is 7 * (4 + 2)?")],
break_point=agent_breakpoint
)
You can now blend text and image capabilities across generation, indexing, and retrieval in Haystack.
New ImageContent Dataclass: A dedicated structure to store image data along with base64_image, mime_type, detail, and metadata.
Image-Aware Chat Generators: Image inputs are now supported in OpenAIChatGenerator
from haystack.dataclasses import ImageContent, ChatMessage
from haystack.components.generators.chat import OpenAIChatGenerator
image_url = "https://cdn.britannica.com/79/191679-050-C7114D2B/Adult-capybara.jpg"
image_content = ImageContent.from_url(image_url)
message = ChatMessage.from_user(
content_parts=["Describe the image in short.", image_content]
)
llm = OpenAIChatGenerator(model="gpt-4o-mini")
print(llm.run([message])["replies"][0].text)
Powerful Multimodal Components:
PDFToImageContent, ImageFileToImageContent, DocumentToImageContent: Convert PDFs, image files, and Documents into ImageContent objects.LLMDocumentContentExtractor: Extract text from images using a vision-enabled LLM.SentenceTransformersDocumentImageEmbedder: Generate embeddings from image-based documents using models like CLIP.DocumentLengthRouter: Route documents based on textual content length—ideal for distinguishing scanned PDFs from text-based ones.DocumentTypeRouter: Route documents automatically based on MIME type metadata.Prompt Building with Image Support: The ChatPromptBuilder now supports templates with embedded images, enabling dynamic multimodal prompt creation.
With these additions, you can now build multimodal agents and RAG pipelines that reason over both text and visual content, unlocking richer interactions and retrieval capabilities.
👉 Learn more about multimodality in our Introduction to Multimodal Text Generation.
Add to_dict and from_dict to ByteStream so it is consistent with our other dataclasses in having serialization and deserialization methods.
Add to_dict and from_dict to classes StreamingChunk, ToolCallResult, ToolCall, ComponentInfo, and ToolCallDelta to make it consistent with our other dataclasses in having serialization and deserialization methods.
Added the tool_invoker_kwargs param to Agent so additional kwargs can be passed to the ToolInvoker like max_workers and enable_streaming_callback_passthrough.
ChatPromptBuilder now supports special string templates in addition to a list of ChatMessage objects. This new format is more flexible and allows structured parts like images to be included in the templatized ChatMessage.
from haystack.components.builders import ChatPromptBuilder
from haystack.dataclasses.chat_message import ImageContent
template = """
{% message role="user" %}
Hello! I am {{user_name}}.
What's the difference between the following images?
{% for image in images %}
{{ image | templatize_part }}
{% endfor %}
{% endmessage %}
"""
images=[
ImageContent.from_file_path("apple-fruit.jpg"),
ImageContent.from_file_path("apple-logo.jpg")
]
builder = ChatPromptBuilder(template=template) builder.run(user_name="John", images=images)
Added convenience class methods to the ImageContent dataclass to create ImageContent objects from file paths and URLs.
Added multiple converters to help convert image data between different formats:
DocumentToImageContent: Converts documents sourced from PDF and image files into ImageContents.
ImageFileToImageContent: Converts image files to ImageContent objects.
ImageFileToDocument: Converts image file references into empty Document objects with associated metadata.
PDFToImageContent: Converts PDF files to ImageContent objects.
Chat Messages with the user role can now include images using the new ImageContent dataclass. We've added image support to OpenAIChatGenerator, and plan to support more model providers over time.
Raise a warning when a pipeline can no longer proceed because all remaining components are blocked from running and no expected pipeline outputs have been produced. This scenario can occur legitimately. For example, in pipelines with mutually exclusive branches where some components are intentionally blocked. To help avoid false positives, the check ensures that none of the expected outputs (as defined by Pipeline().outputs()) have been generated during the current run.
Added source_id_meta_field and split_id_meta_field to SentenceWindowRetriever for customizable metadata field names. Added raise_on_missing_meta_fields to control whether a ValueError is raised if any of the documents at runtime are missing the required meta fields (set to True by default). If False, then the documents missing the meta field will be skipped when retrieving their windows, but the original document will still be included in the results.
Add a ComponentInfo dataclass to the haystack.dataclasses module. This dataclass is used to store information about the component. We pass it to StreamingChunk so we can tell from which component a stream is coming from.
Pass the component_info to the StreamingChunk in the OpenAIChatGenerator, AzureOpenAIChatGenerator, HuggingFaceAPIChatGenerator and HuggingFaceLocalChatGenerator.
Added the enable_streaming_callback_passthrough to the ToolInvoker init, run and run_async methods. If set to True the ToolInvoker will try and pass the streaming_callback function to a tool's invoke method only if the tool's invoke method has streaming_callback in its signature.
Added new HuggingFaceTEIRanker component to enable reranking with Text Embeddings Inference (TEI) API. This component supports both self-hosted Text Embeddings Inference services and Hugging Face Inference Endpoints.
Added a raise_on_failure boolean parameter to OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder. If set to True then the component will raise an exception when there is an error with the API request. It is set to False by default to so the previous behavior of logging an exception and continuing is still the default.
ToolInvoker now executes tool_calls in parallel for both sync and async mode.
Add AsyncHFTokenStreamingHandler for async streaming support in HuggingFaceLocalChatGenerator
Updated StreamingChunk to add the fields tool_calls, tool_call_result, index, and start to make it easier to format the stream in a streaming callback.
HuggingFaceAPIGenerator might no longer work with the Hugging Face Inference API. As of July 2025, the Hugging Face Inference API no longer offers generative models that support the text_generation endpoint. Generative models are now only available through providers that support the chat_completion endpoint. As a result, the HuggingFaceAPIGenerator component might not work with the Hugging Face Inference API. It still works with Hugging Face Inference Endpoints and self-hosted TGI instances. To use generative models via Hugging Face Inference API, please use the HuggingFaceAPIChatGenerator component, which supports the chat_completion endpoint.
All parameters of the Pipeline.draw() and Pipeline.show() methods must now be specified as keyword arguments. Example:
pipeline.draw(
path="output.png",
server_url="https://custom-server.com",
params=None,
timeout=30,
super_component_expansion=False
)
The deprecated async_executor parameter has been removed from the ToolInvoker class. Please use the max_workers parameter instead and a ThreadPoolExecutor with these workers will be created automatically for parallel tool invocations.
The deprecated State class has been removed from the haystack.dataclasses module. The State class is now part of the haystack.components.agents module.
Remove the deserialize_value_with_schema_legacy function from the base_serialization module. This function was used to deserialize State objects created with Haystack 2.14.0 or older. Support for the old serialization format is removed in Haystack 2.16.0.
Add guess_mime_type parameter to Bytestream.from_file_path()
Add the init parameter skip_empty_documents to the DocumentSplitter component. The default value is True. Setting it to False can be useful when downstream components in the Pipeline (like LLMDocumentContentExtractor) can extract text from non-textual documents.
Test that our type validation and connection validation works with builtin python types introduced in 3.9. We found that these types were already supported, we just now add explicit tests for them.
We relaxed the requirement that in ToolCallDelta (introduced in Haystack 2.15) which required the parameters arguments or name to be populated to be able to create a ToolCallDelta dataclass. We remove this requirement to be more in line with OpenAI's SDK and since this was causing errors for some hosted versions of open source models following OpenAI's SDK specification.
Added return_embedding parameter inside InMemoryDocumentStore::init method.
Updated methods bm25_retrieval, and filter_documents to use self.return_embedding to determine whether embeddings are returned.
Updated tests (test_in_memory & test_in_memory_embedding_retriever) to reflect the changes in the InMemoryDocumentStore.
Made doc-parser a core dependency since ComponentTool that uses it is one of the core Tool components.
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. This method is useful for checking that all required connections in a pipeline have a connection and is automatically called in the run method of Pipeline. It is being exposed as public for users who would like to call this method before runtime to validate the pipeline.
Haystack's core modules are now ["type complete"](https://typing.python.org/en/latest/guides/libraries.html#how-much-of-my-library-needs-types), meaning that all function parameters and return types are explicitly annotated. This increases the usefulness of the newly added py.typed marker and sidesteps differences in type inference between the various type checker implementations.
Refactors the HuggingFaceAPIChatGenerator to use the util method _convert_streaming_chunks_to_chat_message. This is to help with being consistent for how we convert StreamingChunks into a final ChatMessage.
We also add ComponentInfo to the StreamingChunks made in HuggingFaceGenerator, and HugginFaceLocalGenerator so we can tell from which component a stream is coming from.
If only system messages are provided as input a warning will be logged to the user indicating that this likely not intended and that they should probably also provide user messages.
Fix _convert_streaming_chunks_to_chat_message which is used to convert Haystack StreamingChunks into a Haystack ChatMessage. This fixes the scenario where one StreamingChunk contains two ToolCallDeltas in StreamingChunk.tool_calls. With this fix this correctly saves both ToolCallDeltas whereas before they were overwriting each other. This only occurs with some LLM providers like Mistral (and not OpenAI) due to how the provider returns tool calls.
Fixed a bug in the print_streaming_chunk utility function that prevented tool call name from being printed.
Fix component_invoker used by ComponentTool to work when a dataclass like ChatMessage is directly passed to component_tool.invoke(...). Previously this would either cause an error or silently skip your input.
RecursiveDocumentSplitter now generates a unique Document.id for every chunk. The meta fields (split_id, parent_id, etc.) are populated [before]() Document creation, so the hash used for id generation is always unique.
In ConditionalRouter fixed the to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. This occurs when a user specifies a route in ConditionalRouter to have multiple outputs.
When calling set_output_types we now also check that the decorator @component.output_types is not present on the run_async method of a Component. Previously we only checked that the Component.run method did not possess the decorator.
@Amnah199 @RafaelJohn9 @anakin87 @bilgeyucel @davidsbatista @julian-risch @kanenorman @kr1shnasomani @mathislucka @mpangrazzi @sjrl @srishti-git1110
Nothing published for this version
We’ve relaxed the requirements for the ToolCallDelta dataclass (introduced in Haystack 2.15). Previously, creating a ToolCallDelta instance required e
ToolCallDelta dataclass (introduced in Haystack 2.15). Previously, creating a ToolCallDelta instance required either the parameters argument or the name to be set. This constraint has now been removed to align more closely with OpenAI's SDK behavior.
The change was necessary as the stricter requirement was causing errors in certain hosted versions of open-source models that adhere to the OpenAI SDK specification.print_streaming_chunk utility function that prevented ToolCall name from being printed.Nothing published for this version
Fix _convert_streaming_chunks_to_chat_message which is used to convert Haystack StreamingChunks into a Haystack ChatMessage. This fixes the scenario w
_convert_streaming_chunks_to_chat_message which is used to convert Haystack StreamingChunks into a Haystack ChatMessage. This fixes the scenario where one StreamingChunk contains two ToolCallDetlas in StreamingChunk.tool_calls. With this fix this correctly saves both ToolCallDeltas whereas before they were overwriting each other. This only occurs with some LLM providers like Mistral (and not OpenAI) due to how the provider returns tool calls.Nothing published for this version
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. T…
ToolInvoker now processes all tool calls passed to run or run_async in parallel using an internal ThreadPoolExecutor. This improves performance by reducing the time spent on sequential tool invocations.ToolInvoker to batch and process multiple tool calls concurrently, allowing Agents to run complex pipelines efficiently with decreased latency.async_executor. ToolInvoker manages its own executor, configurable via the max_workers parameter in init.The new LLMMessagesRouter component that classifies and routes incoming ChatMessage objects to different connections using a generative LLM. This component can be used with general-purpose LLMs and with specialized LLMs for moderation like Llama Guard.
Usage example:
from haystack.components.generators.chat import HuggingFaceAPIChatGenerator
from haystack.components.routers.llm_messages_router import LLMMessagesRouter
from haystack.dataclasses import ChatMessage
chat_generator = HuggingFaceAPIChatGenerator(api_type="serverless_inference_api", api_params={"model": "meta-llama/Llama-Guard-4-12B", "provider": "groq"}, )
router = LLMMessagesRouter(chat_generator=chat_generator, output_names=["unsafe", "safe"], output_patterns=["unsafe", "safe"])
print(router.run([ChatMessage.from_user("How to rob a bank?")]))
HuggingFaceTEIRanker enables end-to-end reranking via the Text Embeddings Inference (TEI) API. It supports both self-hosted TEI services and Hugging Face Inference Endpoints, giving you flexible, high-quality reranking out of the box.
Added a ComponentInfo dataclass to haystack to store information about the component. We pass it to StreamingChunk so we can tell from which component a stream is coming.
Pass the component_info to the StreamingChunk in the OpenAIChatGenerator, AzureOpenAIChatGenerator, HuggingFaceAPIChatGenerator, HuggingFaceGenerator, HugginFaceLocalGenerator and HuggingFaceLocalChatGenerator.
Added the enable_streaming_callback_passthrough to the init, run and run_async methods of ToolInvoker. If set to True the ToolInvoker will try and pass the streaming_callback function to a tool's invoke method only if the tool's invoke method has streaming_callback in its signature.
Added dedicated finish_reason field to StreamingChunk class to improve type safety and enable sophisticated streaming UI logic. The field uses a FinishReason type alias with standard values: "stop", "length", "tool_calls", "content_filter", plus Haystack-specific value "tool_call_results" (used by ToolInvoker to indicate tool execution completion).
Updated ToolInvoker component to use the new finish_reason field when streaming tool results. The component now sets finish_reason="tool_call_results" in the final streaming chunk to indicate that tool execution has completed, while maintaining backward compatibility by also setting the value in meta["finish_reason"].
Added a raise_on_failure boolean parameter to OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder. If set to True then the component will raise an exception when there is an error with the API request. It is set to False by default so the previous behavior of logging an exception and continuing is still the default.
Add AsyncHFTokenStreamingHandler for async streaming support in HuggingFaceLocalChatGenerator
For HuggingFaceAPIGenerator and HuggingFaceAPIChatGenerator all additional key, value pairs passed in api_params are now passed to the initializations of the underlying Inference Clients. This allows passing of additional parameters to the clients like timeout, headers, provider, etc. This means we now can easily specify a different inference provider by passing the provider key in api_params.
Updated StreamingChunk to add the fields tool_calls, tool_call_result, index, and start to make it easier to format the stream in a streaming callback.
ToolCallDelta for the StreamingChunk.tool_calls field to reflect that the arguments can be a string delta.print_streaming_chunk and _convert_streaming_chunks_to_chat_message utility methods to use these new fields. This especially improves the formatting when using print_streaming_chunk with Agent.OpenAIGenerator, OpenAIChatGenerator, HuggingFaceAPIGenerator, HuggingFaceAPIChatGenerator, HuggingFaceLocalGenerator and HuggingFaceLocalChatGenerator to follow the new dataclasses.ToolInvoker to follow the StreamingChunk dataclass.Added a new deserialize_component_inplace function to handle generic component deserialization that works with any component type.
Made doc-parser a core dependency since ComponentTool that uses it is one of the core Tool components.
Make the PipelineBase().validate_input method public so users can use it with the confidence that it won't receive breaking changes without warning. This method is useful for checking that all required connections in a pipeline have a connection and is automatically called in the run method of Pipeline. It is being exposed as public for users who would like to call this method before runtime to validate the pipeline.
For component run Datadog tracing, set the span resource name to the component name instead of the operation name.
Added a trust_remote_code parameter to the SentenceTransformersSimilarityRanker component. When set to True, this enables execution of custom models and scripts hosted on the Hugging Face Hub.
Add a new parameter require_tool_call_ids to ChatMessage.to_openai_dict_format. The default is True, for compatibility with OpenAI's Chat API: if the id field is missing in a Tool Call, an error is raised. Using False is useful for shallow OpenAI-compatible APIs, where the id field is not required.
Haystack's core modules are now "type complete", meaning that all function parameters and return types are explicitly annotated. This increases the usefulness of the newly added py.typed marker and sidesteps differences in type inference between the various type checker implementations.
HuggingFaceAPIChatGenerator now uses the util method _convert_streaming_chunks_to_chat_message. This is to help with being consistent for how we convert StreamingChunks into a final ChatMessage.
async_executor parameter in ToolInvoker is deprecated in favor of max_workers parameter and will be removed in Haystack 2.16.0. You can use max_workers parameter to control the number of threads used for parallel tool calling.to_dict and from_dict of ToolInvoker to properly serialize the streaming_callback init parameter.raise_on_failure=False and an error occurs mid-batch that the following embeddings would be paired with the wrong documents.ComponentTool to work when a dataclass like ChatMessage is directly passed to component_tool.invoke(...). Previously this would either cause an error or silently skip your input.LLMMetadataExtractor that occurred when processing Document objects with None or empty string content. The component now gracefully handles these cases by marking such documents as failed and providing an appropriate error message in their metadata, without attempting an LLM call.Document.id for every chunk. The meta fields (split_id, parent_id, etc.) are populated before Document creation, so the hash used for id generation is always unique.ConditionalRouter fixed the to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. This occurs when a user specifies a route in ConditionalRouter to have multiple outputs.GeneratedAnswer when ChatMessage objects are nested in meta.ComponentTool and Tool when specifying outputs_to_string. Previously an error occurred on deserialization right after serializing if outputs_to_string is not None.set_output_types we now also check that the decorator @component.output_types is not present on the run_async method of a Component. Previously we only checked that the Component.run method did not possess the decorator.is not with != when checking the type List[ChatMessage]. This prevents false mismatches due to Python's is operator comparing object identity instead of equality.__init__.py files. This ensures that short imports like from haystack.components.builders import ChatPromptBuilder work equivalently to from haystack.components.builders.chat_prompt_builder import ChatPromptBuilder, without causing errors or warnings in mypy/Pylance.SuperComponent class can now correctly serialize and deserialize a SuperComponent based on an async pipeline. Previously, the SuperComponent class always assumed the underlying pipeline was synchronous.OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder where if an OpenAI API error occurred mid-batch then the following embeddings would be paired with the wrong documents.Nothing published for this version
In ConditionalRouter fixed the to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. Thi
to_dict and from_dict methods to properly handle the case when output_type is a List of types or a List of strings. This occurs when a user specifies a route in ConditionalRouter to have multiple outputs.outputs_to_string. Previously an error occurred on deserialization right after serializing if outputs_to_string is not None.Nothing published for this version
Fixed a bug in OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder where if an OpenAI API error occurred mid-batch then the following embeddings wo
OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder where if an OpenAI API error occurred mid-batch then the following embeddings would be paired with the wrong documents.raise_on_failure boolean parameter to OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder. If set to True then the component will raise an exception when there is an error with the API request. It is set to False by default so the previous behavior of logging an exception and continuing is still the default.Nothing published for this version
Fixed a mypy issue in the OpenAIChatGenerator and its handling of stream responses. This issue only occurs with mypy \>=1.16.0.
Nothing published for this version
This component replaces the legacy TransformersSimilarityRanker component, which may be deprecated in a future release, with removal following a depre…
We've improved agent workflows with better message handling and streaming support. Agent component now returns a last_message output for quick access to the final message, and can use a streaming_callback to emit tool results in real time. You can use the updated print_streaming_chunk or write your own callback function to enable ToolCall details during streaming.
from haystack.components.websearch import SerperDevWebSearch
from haystack.components.agents import Agent
from haystack.components.generators.utils import print_streaming_chunk
from haystack.tools import tool, ComponentTool
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
web_search = ComponentTool(name="web_search", component=SerperDevWebSearch(top_k=5))
wiki_search = ComponentTool(name="wiki_search", component=SerperDevWebSearch(top_k=5, allowed_domains=["https://www.wikipedia.org/"]))
research_agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-4o-mini"),
system_prompt="""
You are a research agent that can find information on web or specifically on wikipedia.
Use wiki_search tool if you need facts and use web_search tool for latest news on topics.
Use one tool at a time, use the other tool if the retrieved information is not enough.
Summarize the retrieved information before returning response to the user.
""",
tools=[web_search, wiki_search],
streaming_callback=print_streaming_chunk
)
result = research_agent.run(messages=[ChatMessage.from_user("Can you tell me about Florence Nightingale's life?")])
Enabling streaming with print_streaming_chunk function looks like this:
[TOOL CALL]
Tool: wiki_search
Arguments: {"query":"Florence Nightingale"}
[TOOL RESULT]
{'documents': [{'title': 'List of schools in Nottinghamshire', 'link': 'https://www.wikipedia.org/wiki/List_of_schools_in_Nottinghamshire', 'position': 1, 'id': 'a6d0fe00f1e0cd06324f80fb926ba647878fb7bee8182de59a932500aeb54a5b', 'content': 'The Florence Nightingale Academy, Eastwood; The Flying High Academy, Mansfield; Forest Glade Primary School, Sutton-in-Ashfield; Forest Town Primary School ...', 'blob': None, 'score': None, 'embedding': None, 'sparse_embedding': None}], 'links': ['https://www.wikipedia.org/wiki/List_of_schools_in_Nottinghamshire']}
...
Print the last_message
print("Final Answer:", result["last_message"].text)
>>> Final Answer: Florence Nightingale (1820-1910) was a pioneering figure in nursing and is often hailed as the founder of modern nursing. She was born...
Additionally, AnswerBuilder stores all generated messages in all_messages meta field of GeneratedAnswer and supports a new last_message_only mode for lightweight flows where only the final message needs to be processed.
We extended pipeline.draw() and pipeline.show(), which save pipeline diagrams to images files or display them in Jupyter notebooks. You can now pass super_component_expansion=True to expand any SuperComponents and draw more detailed visualizations.
Here is an example with a pipeline containing MultiFileConverter and DocumentPreprocssor SuperComponents. After installing the dependencies that the MultiFileConverter needs for all supported file formats via pip install haystack-ai pypdf markdown-it-py mdit_plain trafilatura python-pptx python-docx jq openpyxl tabulate pandas, you can run:
from pathlib import Path
from haystack import Pipeline
from haystack.components.converters import MultiFileConverter
from haystack.components.preprocessors import DocumentPreprocessor
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore
document_store = InMemoryDocumentStore()
pipeline = Pipeline()
pipeline.add_component("converter", MultiFileConverter())
pipeline.add_component("preprocessor", DocumentPreprocessor())
pipeline.add_component("writer", DocumentWriter(document_store = document_store))
pipeline.connect("converter", "preprocessor")
pipeline.connect("preprocessor", "writer")
# expanded pipeline that shows all components
path = Path("expanded_pipeline.png")
pipeline.draw(path=path, super_component_expansion=True)
# original pipeline
path = Path("original_pipeline.png")
pipeline.draw(path=path)
We added a new SentenceTransformersSimilarityRanker component that uses the Sentence Transformers library to rank documents based on their semantic similarity to the query. This component replaces the legacy TransformersSimilarityRanker component, which may be deprecated in a future release, with removal following a deprecation period. The SentenceTransformersSimilarityRanker also allows choosing different inference backends: PyTorch, ONNX, and OpenVINO. For example, after installing sentence-transformers>=4.1.0, you can run:
from haystack.components.rankers import SentenceTransformersSimilarityRanker
from haystack.utils.device import ComponentDevice
onnx_ranker = SentenceTransformersSimilarityRanker(
model="sentence-transformers/all-MiniLM-L6-v2",
token=None,
device=ComponentDevice.from_str("cpu"),
backend="onnx",
)
onnx_ranker.warm_up()
docs = [Document(content="Berlin"), Document(content="Sarajevo")]
output = onnx_ranker.run(query="City in Germany", documents=docs)
ranked_docs = output["documents"]
py.typed file to Haystack to enable type information to be used by downstream projects, in line with PEP 561. This means Haystack's type hints will now be visible to type checkers in projects that depend on it. Haystack is primarily type checked using mypy (not pyright) and, despite our efforts, some type information can be incomplete or unreliable. If you use static type checking in your own project, you may notice some changes: previously, Haystack's types were effectively treated as Any, but now actual type information will be available and enforced. We'll continue improving typing with the next release.deserialize_tools_inplace utility function has been removed. Use deserialize_tools_or_toolset_inplace instead, importing it as follows: from haystack.tools import deserialize_tools_or_toolset_inplace.Added run_async method to ToolInvoker class to allow asynchronous tool invocations.
Agent can now stream tool result with run_async method as well.
Introduced serialize_value and deserialize_value utility methods for consistent value (de)serialization across modules.
Moved the State class to the agents.state module and added serialization and deserialization capabilities.
Add support for multiple outputs in ConditionalRouter
Implement JSON-safe serialization for OpenAI usage data by converting token counts and details (like CompletionTokensDetails and PromptTokensDetails) into plain dictionaries.
Added a new SentenceTransformersSimilarityRanker component that uses the Sentence Transformers library to rank documents based on their semantic similarity to the query. This component is a replacement for the legacy TransformersSimilarityRanker component, which may be deprecated in a future release, with removal following after a deprecation period. The SentenceTransformersSimilarityRanker also allows choosing different inference backends: PyTorch, ONNX, and OpenVINO. To use the SentenceTransformersSimilarityRanker, you need to install sentence-transformers>=4.1.0.
Add a streaming_callback parameter to ToolInvoker to enable streaming of tool results. Note that tool_result is emitted only after the tool execution completes and is not streamed incrementally.
Update print_streaming_chunk to print ToolCall information if it is present in the chunk's metadata.
Update Agent to forward the streaming_callback to ToolInvoker to emit tool results during tool invocation.
Enhance SuperComponent's type compatibility check to return the detected common type between two input types.
When using HuggingFaceAPIChatGenerator with streaming, the returned ChatMessage now contains the number of prompt tokens and completion tokens in its meta data. Internally, the HuggingFaceAPIChatGenerator requests an additional streaming chunk that contains usage data. It then processes the usage streaming chunk to add usage meta data to the returned ChatMessage.
We now have a Protocol for TextEmbedder. The protocol makes it easier to create custom components or SuperComponents that expect any TextEmbedder as init parameter.
We added a Component signature validation method that details the mismatches between the run and run_async method signatures. This allows a user to debug custom components easily.
Enhanced the AnswerBuilder component with two agent-friendly features:
meta field of the GeneratedAnswer objects under an all_messages key, improving traceability and debugging capabilities.last_message_only parameter that, when set to True, processes only the last message in the replies while still preserving the complete conversation history in metadata. This is particularly useful for agent workflows where only the final response needs to be processed.A variety of improvements have been made so an Agent component can be directly used in ComponentTool enabling straightforward building of Multi-Agent systems. These improvements include:
last_message field to the Agent's output which returns the last generated ChatMessage._default_output_handler in the ToolInvoker to try and first serialize the outputs in the tool result before converting it into a string. This is especially relevant for getting a better representation when stringifying dataclasses like ChatMessage.Added type hints to the component decorator. This improves support for Pyright/Pylance, enabling IDEs like VSCode to show docstrings for components.
Updated pipeline execution logic to use a new utility method _deepcopy_with_exceptions, which attempts to deep copy an object and safely falls back to the original object if copying fails. Additionally _deepcopy_with_exceptions skips deep-copying of Component, Tool, and Toolset instances when used as runtime parameters. This prevents errors and unintended behavior caused by trying to deepcopy objects that contain non-copyable attributes (e.g. Jinja2 templates, clients). Previously, standard deepcopy was used on inputs and outputs which occasionally lead to errors since certain Python objects cannot be deepcopied.
Refactored JSON Schema generation for ComponentTool parameters using Pydantic’s model_json_schema, enabling expanded type support (e.g. Union, Enum, Dict, etc.). We also convert dataclasses to Pydantic models before calling model_json_schema to preserve docstring descriptions of the parameters in the schema. This means dataclasses like ChatMessage, Document, etc. now have correctly defined JSON schemas.
The draw() and show() methods from Pipeline now have an extra boolean parameter, super_component_expansion, which, when set to True and if the pipeline contains SuperComponents, the visualisation diagram will show the internal structure of super-components as if they were components part of the pipeline instead of a "black-box" with the name of the SuperComponent.
Improve the type annotations for @component and the Component protocol. The type checker can now ensure that a @component class provides a compatible run() method, whose required return type has been changed from Dict[str, Any] (invariant) to the Mapping[str, Any] to allow TypedDict to be used for output types.
ComponentTool now preserves and combines docstrings from underlying pipeline components when wrapping a SuperComponent. When a SuperComponent is used with ComponentTool, two key improvements are made:
These changes make SuperComponents much more useful with LLM function calling as the LLM will get detailed information about both the component's purpose and its parameters.
Adds local_files_only parameter to SentenceTransformersDocumentEmbedder and SentenceTransformersTextEmbedder to allow loading models in offline mode.
The DocumentRecallEvaluator was updated. Now, when in MULTI_HIT mode, the division is over the unique ground truth documents instead of the total number of ground truth documents. We also added checks for emptiness. If there are no retrieved documents or all of them have an empty string as content, we return 0.0 and log a warning. Likewise, if there are no ground truth documents or all of them have an empty string as content, we return 0.0 and log a warning.
State class in the dataclasses module. Users are encouraged to transition to the new version of State now located in the agents.state module. A deprecation warning has been added to guide this migration.__deepcopy__ of ComponentTool to gracefully handle NotImplementedError when trying to deepcopy attributes.RecursiveDocumentSplitter was fixed for the case where a split_text is longer than the split_length and recursive chunking is triggered.huggingface_hub>=0.31.0. In the huggingface_hub library, arguments attribute of ChatCompletionInputFunctionDefinition has been renamed to parameters. Our implementation is compatible with both the legacy version and the new one.HuggingFaceAPIChatGenerator now checks the type of the arguments variable in the tool calls returned by the Hugging Face API. If arguments is a JSON string, it is parsed into a dictionary. Previously, the arguments type was not checked, which sometimes led to failures later in the tool workflow.ctx.run(...) so we can preserve context like the active tracing span. This now means if your component 1) only has a sync run method and 2) it logs something to the tracer then this trace will be properly nested within the parent context.LLMMetadataExtractor that occurred when processing Document objects with None or empty string content. The component now gracefully handles these cases by marking such documents as failed and providing an appropriate error message in their metadata, without attempting an LLM call.component_tool.invoke(...). Previously this would either cause an error or silently skip your input.Special thanks and congratulations to our first time contributors!
Full Changelog: https://github.com/deepset-ai/haystack/compare/v2.13.0...v2.14.0
Nothing published for this version
Nothing published for this version
Updated pipeline execution logic to use a new utility method \_deepcopy_with_exceptions , which attempts to deep copy an object and safely falls back
Update the \_\_deepcopy\_\_ of ComponentTool to gracefully handle NotImplementedError when trying to deepcopy attributes.
The deprecated api, api_key, and api_params parameters for LLMEvaluator, ContextRelevanceEvaluator, and FaithfulnessEvaluator have been removed. By de…
Haystack's Agent got several improvements!
Agent Tracing
Agent tracing now provides deeper visibility into the agent's execution. For every call, the inputs and outputs of the ChatGenerator and ToolInvoker are captured and logged using dedicated child spans. This makes it easier to debug, monitor, and analyze how an agent operates step-by-step.
Below is an example of what the trace looks like in Langfuse:
<p align="center"> <img width="500" alt="Langfuse UI for tracing" src="https://github.com/user-attachments/assets/ee00e66d-ee6b-4c87-980e-5aa7f948b6af" /> </p>
# pip install langfuse-haystack
from haystack_integrations.components.connectors.langfuse.langfuse_connector import LangfuseConnector
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
tracer = LangfuseConnector("My Haystack Agent")
agent = Agent(
system_prompt="You help provide the weather for cities"
chat_generator=OpenAIChatGenerator(),
tools=[weather_tool],
)
Async Support
Additionally, there's a new run_async method to enable built-in async support for Agent. Just use run_async instead of the run method. Here's an example of an async web search agent:
# set `SERPERDEV_API_KEY` and `OPENAI_API_KEY` as env variables
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.components.websearch import SerperDevWebSearch
from haystack.dataclasses import ChatMessage
from haystack.tools.component_tool import ComponentTool
web_tool = ComponentTool(component=SerperDevWebSearch())
web_search_agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=[web_tool],
)
result = await web_search_agent.run_async(
messages=[ChatMessage.from_user("Find information about Haystack by deepset")]
)
The new Toolset groups multiple Tool instances into a single manageable unit. It simplifies the passing of tools to components like ChatGenerator, ToolInvoker, or Agent, and supports filtering, serialization, and reuse.
Check out the MCPToolset for dynamic tool discovery from an MCP server.
from haystack.tools import Toolset
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
math_toolset = Toolset([tool_one, tool_two, ...])
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-4o-mini"),
tools=math_toolset
)
Creating a custom SuperComponents just got even simpler. Now, all you need to do is define a class with a pipeline attribute and decorate it with @super_component. Haystack takes care of the rest!
Here's an example of building a custom HybridRetriever using the @super_component decorator:
# pip install haystack-ai datasets "sentence-transformers>=3.0.0"
from haystack import Document, Pipeline, super_component
from haystack.components.joiners import DocumentJoiner
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers import InMemoryBM25Retriever, InMemoryEmbeddingRetriever
from haystack.document_stores.in_memory import InMemoryDocumentStore
from datasets import load_dataset
@super_component
class HybridRetriever:
def __init__(self, document_store: InMemoryDocumentStore, embedder_model: str = "BAAI/bge-small-en-v1.5"):
embedding_retriever = InMemoryEmbeddingRetriever(document_store)
bm25_retriever = InMemoryBM25Retriever(document_store)
text_embedder = SentenceTransformersTextEmbedder(embedder_model)
document_joiner = DocumentJoiner(join_mode="reciprocal_rank_fusion")
self.pipeline = Pipeline()
self.pipeline.add_component("text_embedder", text_embedder)
self.pipeline.add_component("embedding_retriever", embedding_retriever)
self.pipeline.add_component("bm25_retriever", bm25_retriever)
self.pipeline.add_component("document_joiner", document_joiner)
self.pipeline.connect("text_embedder", "embedding_retriever")
self.pipeline.connect("bm25_retriever", "document_joiner")
self.pipeline.connect("embedding_retriever", "document_joiner")
dataset = load_dataset("HaystackBot/medrag-pubmed-chunk-with-embeddings", split="train")
docs = [Document(content=doc["contents"], embedding=doc["embedding"]) for doc in dataset]
document_store = InMemoryDocumentStore()
document_store.write_documents(docs)
query = "What treatments are available for chronic bronchitis?"
result = HybridRetriever(document_store).run(text=query, query=query)
print(result)
New ready-made SuperComponents: MultiFileConverter, DocumentPreprocessor
There are also two ready-made SuperComponents, MultiFileConverter and DocumentPreprocessor, that encapsulate widely used common logic for indexing pipelines.
📚 Learn more about SuperComponents and get the full code example in the Tutorial: Creating Custom SuperComponents
api, api_key, and api_params parameters for LLMEvaluator, ContextRelevanceEvaluator, and FaithfulnessEvaluator have been removed. By default, these components will continue to use OpenAI in JSON mode. To customize the LLM, use the chat_generator parameter with a ChatGenerator instance configured to return a response in JSON format. For example:chat_generator=OpenAIChatGenerator(generation_kwargs={"response_format": {"type": "json_object"}})
generator_api and generator_api_params initialization parameters of LLMMetadataExtractor and the LLMProvider enum have been removed. Use chat_generator instead to configure the underlying LLM. In order for the component to work, the LLM should be configured to return a JSON object. For example, if using OpenAI, you should initialize the LLMMetadataExtractor withchat_generator=OpenAIChatGenerator(generation_kwargs={"response_format": {"type": "json_object"}})
OpenAITextEmbedder.run_async method to HuggingFaceAPIDocumentEmbedder. This method enriches Documents with embeddings. It supports the same parameters as the run method. It returns a coroutine that can be awaited.http_client_kwargs (proxy, SSL) for:
AzureOpenAIGenerator, OpenAIGenerator and DALLEImageGeneratorOpenAIDocumentEmbedder and OpenAITextEmbedderRemoteWhisperTranscriberOpenAIChatGenerator and AzureOpenAIChatGenerator now support custom HTTP client config via http_client_kwargs, enabling proxy and SSL setup.HuggingFaceAPITextEmbedder now also has support for a run() method in an asynchronous way, i.e., run_async.output_mapping.AzureOpenAITextEmbedder and AzureOpenAIDocumentEmbedder now support custom HTTP client config via http_client_kwargs, enabling proxy and SSL setup.AzureOpenAIDocumentEmbedder component now inherits from the OpenAIDocumentEmbedder component, enabling asynchronous usage.AzureOpenAITextEmbedder component now inherits from the OpenAITextEmbedder component, enabling asynchronous usage.OpenAIDocumentEmbedder component.component_name and component_type attributes to PipelineRuntimeError.
PipelineRuntimeErrorChatGenerator Protocol no longer requires to_dict and from_dict methods.deserialize_tools_inplace has been deprecated and will be removed in Haystack 2.14.0. Use deserialize_tools_or_toolset_inplace instead.OpenAITextEmbedder no longer replaces newlines with spaces in the text to embed. This was only required for the discontinued v1 embedding models.OpenAIDocumentEmbedder and AzureOpenAIDocumentEmbedder no longer replace newlines with spaces in the text to embed. This was only required for the discontinued v1 embedding models.ChatMessage.from_dict to handle cases where optional fields like name and meta are missing.AsyncPipeline, the span tag name is updated from hasytack.component.outputs to haystack.component.output. This matches the tag name used in Pipeline and is the tag name expected by our tracers.Nothing published for this version
Nothing published for this version
Fix ChatMessage.from_dict to handle cases where optional fields like name and meta are missing.
In Agent we make sure state_schema is always initialized to have 'messages'. Previously this was only happening at run time which is why pipeline.conn
Updated SentenceTransformersDiversityRanker to use the token parameter internally instead of the deprecated use_auth_token. The public API of this com…
The Agent component enables tool-calling functionality with provider-agnostic chat model support and can be used as a standalone component or within a pipeline.
With SERPERDEV_API_KEY and OPENAI_API_KEY defined, a Web Search Agent is as simple as:
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.components.websearch import SerperDevWebSearch
from haystack.dataclasses import ChatMessage
from haystack.tools.component_tool import ComponentTool
web_tool = ComponentTool(
component=SerperDevWebSearch(),
)
agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=[web_tool],
)
result = agent.run(
messages=[ChatMessage.from_user("Find information about Haystack by deepset")]
)
The Agent supports streaming responses, customizable exit conditions, and a flexible state management system that enables tools to share and modify data during execution:
agent = Agent(
chat_generator=OpenAIChatGenerator(),
tools=[web_tool, weather_tool],
exit_conditions=["text", "weather_tool"],
state_schema = {...},
streaming_callback=streaming_callback,
)
SuperComponent allows you to wrap complex pipelines into reusable components. This makes it easy to reuse them across your applications. Just initialize a SuperComponent with a pipeline:
from haystack import Pipeline, SuperComponent
with open("pipeline.yaml", "r") as file:
pipeline = Pipeline.load(file)
super_component = SuperComponent(pipeline)
That's not all! To show the benefits, there are three ready-made SuperComponents in haystack-experimental.
For example, there is a MultiFileConverter that wraps a pipeline with converters for CSV, DOCX, HTML, JSON, MD, PPTX, PDF, TXT, and XSLX. After installing the integration dependencies pip install pypdf markdown-it-py mdit_plain trafilatura python-pptx python-docx jq openpyxl tabulate pandas, you can run with any of the supported file types as input:
from haystack_experimental.super_components.converters import MultiFileConverter
converter = MultiFileConverter()
converter.run(sources=["test.txt", "test.pdf"], meta={})
Here's an example of creating a custom SuperComponent from any Haystack pipeline:
from haystack import Pipeline, SuperComponent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.components.builders import ChatPromptBuilder
from haystack.components.retrievers import InMemoryBM25Retriever
from haystack.dataclasses.chat_message import ChatMessage
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.dataclasses import Document
document_store = InMemoryDocumentStore()
documents = [
Document(content="Paris is the capital of France."),
Document(content="London is the capital of England."),
]
document_store.write_documents(documents)
prompt_template = [
ChatMessage.from_user(
'''
According to the following documents:
{% for document in documents %}
{{document.content}}
{% endfor %}
Answer the given question: {{query}}
Answer:
'''
)
]
prompt_builder = ChatPromptBuilder(template=prompt_template, required_variables="*")
pipeline = Pipeline()
pipeline.add_component("retriever", InMemoryBM25Retriever(document_store=document_store))
pipeline.add_component("prompt_builder", prompt_builder)
pipeline.add_component("llm", OpenAIChatGenerator())
pipeline.connect("retriever.documents", "prompt_builder.documents")
pipeline.connect("prompt_builder.prompt", "llm.messages")
# Create a super component with simplified input/output mapping
wrapper = SuperComponent(
pipeline=pipeline,
input_mapping={
"query": ["retriever.query", "prompt_builder.query"],
},
output_mapping={"llm.replies": "replies"}
)
# Run the pipeline with simplified interface
result = wrapper.run(query="What is the capital of France?")
print(result)
# {'replies': [ChatMessage(_role=<ChatRole.ASSISTANT: 'assistant'>,
# _content=[TextContent(text='The capital of France is Paris.')],...)
Updated ChatMessage serialization and deserialization. ChatMessage.to_dict() now returns a dictionary with the keys: role, content, meta, and name. ChatMessage.from_dict() supports this format and maintains compatibility with older formats.
If your application consumes the result of ChatMessage.to_dict(), update your code to handle the new format. No changes are needed if you're using ChatPromptBuilder in a Pipeline.
LLMEvaluator, ContextRelevanceEvaluator, and FaithfulnessEvaluator now internally use a ChatGenerator instance instead of a Generator instance. The public attribute generator has been replaced with _chat_generator.
to_pandas, comparative_individual_scores_report and score_report were removed from EvaluationRunResult, please use detailed_report, comparative_detailed_report and aggregated_report instead.
outputs_to_string to Tool and ComponentTool to allow users to customize how the output of a Tool should be converted into a string so that it can be provided back to the ChatGenerator in a ChatMessage. If outputs_to_string is not provided, a default converter is used within ToolInvoker. The default handler uses the current default behavior.split_mode to the CSVDocumentSplitter component to control the splitting mode. The new parameter can be set to row-wise to split the CSV file by rows. The default value is threshold, which is the previous behavior.run_async method to HuggingFaceLocalChatGenerator. This method internally uses ThreadPoolExecutor to return coroutines that can be awaited.Documents. It accepts a link_format parameter that can be set to "markdown" or "plain". By default, no hyperlink addresses are extracted as before.azure_ad_token_provider to all Azure OpenAI components: AzureOpenAIGenerator, AzureOpenAIChatGenerator, AzureOpenAITextEmbedder and AzureOpenAIDocumentEmbedder. This parameter optionally accepts a callable that returns a bearer token, enabling authentication via Azure AD.
default_azure_token_provider in haystack/utils/azure.py. This function provides a default token provider that is serializable by Haystack. Users can now pass default_azure_token_provider as the azure_ad_token_provider or implement a custom token provider.ChatPromptBuilder. In the same way as the PromptBuilder, the ChatPromptBuilder now supports arrow to work with datetime.LLMEvaluator, ContextRelevanceEvaluator, and FaithfulnessEvaluator now accept a chat_generator initialization parameter, consisting of ChatGenerator instance pre-configured to return a JSON object. Previously, these components only supported OpenAI and LLMs with OpenAI-compatible APIs. Regardless of whether the evaluator components are initialized with api, api_key, and api_params or the new chat_generator parameter, the serialization format will now only include chat_generator in preparation for the future removal of api, api_key, and api_params.Agent, the messages are stored and accumulated in State. This means:
exit_condition to exit_conditions to reflect that.Agent, we check all messages from the LLM when doing an exit condition check. For example, it's possible the LLM returns multiple messages, such as multiple tool calls, or includes messages with reasoning. Now we check all messages before assessing if we should exit the loop.Agent component checks whether the ChatGenerator it is initialized with supports tools. If it doesn't, the Agent raises a TypeError.BranchJoiner to more understandable and better highlight where it's useful.ChatPromptBuilder and PromptBuilder when prompt variables are present and required_variables is unset to help users avoid unexpected execution in multi-branch pipelines. The warning recommends users to set required_variables.api, api_key, and api_params parameters for LLMEvaluator, ContextRelevanceEvaluator, and FaithfulnessEvaluator are now deprecated and will be removed in Haystack 2.13.0. By default, these components will continue to use OpenAI in JSON mode. To configure a specific LLM, use the chat_generator parameter.chat_generator instead to configure the underlying LLM. For example, change generator_api=LLMProvider.OPENAI to chat_generator=OpenAIChatGenerator().Document.from_dict() in haystack-ai>=2.11.0 could not properly deserialize a Document dictionary obtained with document.to_dict(flatten=False) in haystack-ai<=2.10.0.max_retries initialization parameter is correctly set when equal 0 in AzureOpenAIGenerator, AzureOpenAIChatGenerator, AzureOpenAITextEmbedder and AzureOpenAIDocumentEmbedder.component.output_types decorator. The type hinting for the decorator was originally introduced to avoid overshadowing the type hinting of the run method and allow proper static type checking. This update extends support to asynchronous run_async methods.MistralChatGenerator not returning a finish_reason when using streaming. Fixed by adjusting how we look for the finish_reason when processing streaming chunks. Now, the last non-None finish_reason is used to handle differences between OpenAI and Mistral.Nothing published for this version
Refactored the processing of streaming chunks from OpenAI to simplify logic.
Add dataframe to legacy fields for the Document dataclass. This fixes a bug where Document.from_dict() in haystack-ai\>=2.11.0 could not properly dese
Nothing published for this version
The ExtractedTableAnswer dataclass and the dataframe field in the Document dataclass, deprecated in Haystack 2.10.0, have now been removed. pandas is…
With lazy importing, importing individual components now requires 50% less CPU time on average. Overall import performance has also significantly improved: for example, import haystack now consumes only 2-5% of the CPU time it previously did.
As of this release, all chat generators and retrievers in the core package now include a run_async method, enabling asynchronous execution at the component level. When used in an AsyncPipeline, this method runs automatically, providing native async capabilities.
<p align="center"> <img width="600" alt="AsyncPipeline vs Pipeline" src="https://github.com/user-attachments/assets/9d954472-53ea-4efd-8dc0-7408643f87ae" /> </p>
MSGToDocument ComponentUse MSGToDocument to convert Microsoft Outlook .msg files into Haystack documents. This component extracts the email metadata (such as sender, recipients, CC, BCC, subject) and body content and converts any file attachments into ByteStream objects.
Set connection_type_validation to false when initializing Pipeline to disable type validation for pipeline connections. This will allow you to connect any edges and bypass errors you might get, for example, when you connect Optional[str] output to str input.
The ExtractedTableAnswer dataclass and the dataframe field in the Document dataclass, deprecated in Haystack 2.10.0, have now been removed. pandas is no longer a required dependency for Haystack, making the installation lighter. If a component you use requires pandas, an informative error will be raised, prompting you to install it. For details and motivation, see the GitHub discussion #8688.
Starting from Haystack 2.11.0 Python 3.8 is no longer supported. Python 3.8 reached its end of life on October 2024.
The AzureOCRDocumentConverter no longer produces Document objects with the deprecated dataframe field.
Am I affected?
dataframe field in Document objects generated by AzureOCRDocumentConverter, you are affected.DeprecationWarning in Haystack 2.10 when initializing a Document with a dataframe, this change will now remove that field entirely.How to handle the change:
dataframe, AzureOCRDocumentConverter now represents tables as CSV-formatted text in the content field of the Document.dataframe. If needed, you can convert the CSV text back into a dataframe using pandas.read_csv().run_async method to HuggingFaceAPIChatGenerator. This method relies internally on the AsyncInferenceClient from huggingface to generate chat completions and supports the same parameters as the run method. It returns a coroutine that can be awaited.run_async method to OpenAIChatGenerator. This method internally uses the async version of the OpenAI client to generate chat completions and supports the same parameters as the run method. It returns a coroutine that can be awaited.run_async method to DocumentWriter. This method supports the same parameters as the run method and relies on the DocumentStore to implement write_documents_async. It returns a coroutine that can be awaited.run_async method to AzureOpenAIChatGenerator. This method uses AsyncAzureOpenAI to generate chat completions and supports the same parameters as the run method. It returns a coroutine that can be awaited.run_async method to HuggingFaceLocalChatGenerator. This method internally uses ThreadPoolExecutor to return coroutines that can be awaited.Pipeline.show and Pipeline.draw methods. This allows users to customize the timeout as needed.spacy backend. Additionally, you may encounter issues installing openai-whisper, which is required by the LocalWhisperTranscriber component, if you use uv or poetry for installation. In this case, we recommend using pip for installation.EvaluationRunResult can now output the results in JSON, a pandas Dataframe or in a CSV file.typing. prefix for standard typing library types (e.g., List[str] instead of typing.List[str]).EvaluationRunResult is now optional and the methods score_report, to_pandas and comparative_individual_scores_report are deprecated and will be removed in the next haystack release.ChatMessage.to_openai_dict_format utility method, include the name field in the returned dictionary, if present. Previously, the name field was erroneously skipped.additionalProperties: False in the tool schema when tool_strict is set to True.output_type of a ConditionalRouter was not being serialized correctly. This would cause the router to work incorrectly after being serialized and deserialized.typing.Any when using serialize_type utilitydescription anymore.haystack/utils/type_serialization.py to handle Optional types correctly.Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixed accumulation of a tools arguments when streaming with an OpenAIChatGenerator
Nothing published for this version
Pipelines with components that return plain pandas dataframes failed. The comparison of socket values is now 'is not' instead of '!=' to avoid errors
Nothing published for this version
ComponentTool does not truncate 'description' anymore.
Nothing published for this version
Removed the deprecated NLTKDocumentSplitter, it's functionalities are now supported by the `DocumentSplitter`.
Pipeline.run() LogicThe new Pipeline.run() logic fixes common pipeline issues, including exceptions, incorrect component execution, missing intermediate outputs, and premature execution of lazy variadic components. While most pipelines should remain unaffected, we recommend carefully reviewing your pipeline executions if you are using cyclic pipelines or pipelines with lazy variadic components to ensure their behavior has not changed. You can use this tool to compare the execution traces of your pipeline with the old and new logic.
AsyncPipeline for Async ExecutionTogether with the new Pipeline.run logic, AsyncPipeline enables asynchronous execution, allowing pipeline components to run concurrently whenever possible. This leads to significant speed improvements, especially for pipelines processing data in parallel branches such as hybrid retrieval setting.
<p align="center"> <img width="600" alt="AsyncPipeline vs Pipeline" src="https://github.com/user-attachments/assets/9d954472-53ea-4efd-8dc0-7408643f87ae" /> </p> <details> <summary><h4>Source Codes</h4></summary>
Hybrid Retrieval
hybrid_rag_retrieval = AsyncPipeline()
hybrid_rag_retrieval.add_component("text_embedder", SentenceTransformersTextEmbedder())
hybrid_rag_retrieval.add_component("embedding_retriever", InMemoryEmbeddingRetriever(document_store=document_store))
hybrid_rag_retrieval.add_component("bm25_retriever", InMemoryBM25Retriever(document_store=document_store))
hybrid_rag_retrieval.connect("text_embedder", "embedding_retriever")
hybrid_rag_retrieval.connect("bm25_retriever", "document_joiner")
hybrid_rag_retrieval.connect("embedding_retriever", "document_joiner")
async def run_inner():
return await hybrid_rag_retrieval.run({
"text_embedder": {"text": query},
"bm25_retriever": {"query": query}
})
results = asyncio.run(run_inner())
Parallel Translation Pipeline
from haystack.components.builders import ChatPromptBuilder
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack import AsyncPipeline
from haystack.utils import Secret
# Create prompt builders with templates at initialization
spanish_prompt_builder = ChatPromptBuilder(template="Translate this message to Spanish: {{user_message}}")
turkish_prompt_builder = ChatPromptBuilder(template="Translate this message to Turkish: {{user_message}}")
thai_prompt_builder = ChatPromptBuilder(template="Translate this message to Thai: {{user_message}}")
# Create LLM instances
spanish_llm = OpenAIChatGenerator()
turkish_llm = OpenAIChatGenerator()
thai_llm = OpenAIChatGenerator()
# Create and configure pipeline
pipe = AsyncPipeline()
# Add components
pipe.add_component("spanish_prompt_builder", spanish_prompt_builder)
pipe.add_component("turkish_prompt_builder", turkish_prompt_builder)
pipe.add_component("thai_prompt_builder", thai_prompt_builder)
pipe.add_component("spanish_llm", spanish_llm)
pipe.add_component("turkish_llm", turkish_llm)
pipe.add_component("thai_llm", thai_llm)
# Connect components
pipe.connect("spanish_prompt_builder.prompt", "spanish_llm.messages")
pipe.connect("turkish_prompt_builder.prompt", "turkish_llm.messages")
pipe.connect("thai_prompt_builder.prompt", "thai_llm.messages")
user_message = """
In computer programming, the async/await pattern is a syntactic feature of many programming languages that
allows an asynchronous, non-blocking function to be structured in a way similar to an ordinary synchronous function.
It is semantically related to the concept of a coroutine and is often implemented using similar techniques,
and is primarily intended to provide opportunities for the program to execute other code while waiting
for a long-running, asynchronous task to complete, usually represented by promises or similar data structures.
"""
# Run the pipeline with simplified input
res = pipe.run(data={"user_message": user_message})
# Print results
print("Spanish translation:", res["spanish_llm"]["generated_messages"][0].text)
print("Turkish translation:", res["turkish_llm"]["generated_messages"][0].text)
print("Thai translation:", res["thai_llm"]["generated_messages"][0].text)
</details>
Tool calling is now universally supported across all chat generators, making it easier than ever for developers to port tools across different platforms. Simply switch the chat generator used, and tooling will work seamlessly without any additional configuration. This update applies across AzureOpenAIChatGenerator, HuggingFaceLocalChatGenerator, and all core integrations, including AnthropicChatGenerator, CohereChatGenerator, AmazonBedrockChatGenerator, and VertexAIGeminiChatGenerator. With this enhancement, tool usage becomes a native capability across the ecosystem, enabling more advanced and interactive agentic applications.
Pipeline visualization is now more flexible, allowing users to render pipeline graphs locally without requiring an internet connection or sending data to an external service. By running a local Mermaid server with Docker, you can generate visual representations of your pipelines using draw() or show(). Learn more in Visualizing Pipelines
This release introduces new components that enhance document processing capabilities. CSVDocumentSplitter and CSVDocumentCleaner make handling CSV files more efficient. LLMMetadaExtractor leverages an LLM to analyze documents and enrich them with relevant metadata, improving searchability and retrieval accuracy.
DOCXToDocument converter now returns a Document object with DOCX metadata stored in the meta field as a dictionary under the key docx. Previously, the metadata was represented as a DOCXMetadata dataclass. This change does not impact reading from or writing to a Document Store.NLTKDocumentSplitter, it's functionalities are now supported by the DocumentSplitter.Added a new component ListJoiner which joins lists of values from different components to a single list.
Introduced the OpenAPIConnector component, enabling direct invocation of REST endpoints as specified in an OpenAPI specification. This component is designed for direct REST endpoint invocation without LLM-generated payloads, users needs to pass the run parameters explicitly. Example:
from haystack.utils import Secret
from haystack.components.connectors.openapi import OpenAPIConnector
connector = OpenAPIConnector(openapi_spec="https://bit.ly/serperdev_openapi", credentials=Secret.from_env_var("SERPERDEV_API_KEY"))
response = connector.run(operation_id="search", parameters={"q": "Who was Nikola Tesla?"} )
Adding a new component, LLMMetadaExtractor, which can be used in an indexing pipeline to extract metadata from documents based on a user-given prompt and return the documents with the metadata field with the output of the LLM.
Introduced CSVDocumentCleaner component for cleaning CSV documents.
Introducing CSVDocumentSplitter: The CSVDocumentSplitter splits CSV documents into structured sub-tables by recursively splitting by empty rows and columns larger than a specified threshold. This is particularly useful when converting Excel files which can often have multiple tables within one sheet.
SentenceTransformersDocumentEmbedder and SentenceTransformersTextEmbedder to accept an additional parameter, which is passed directly to the underlying SentenceTransformer.encode method for greater flexibility in embedding customization.completion_start_time metadata to track time-to-first-token (TTFT) in streaming responses from Hugging Face API and OpenAI (Azure).MetadataRouter:
_parse_date, which first attempts datetime.fromisoformat(value) for backward compatibility and then falls back to dateutil.parser.parse() for broader ISO 8601 support._ensure_both_dates_naive_or_aware, which ensures both datetimes are either naive or aware. If one is missing a timezone, it is assigned the timezone of the other for consistency.Pipeline.from_dict receives an invalid type (e.g. empty string), an informative PipelineError is now raised.CSVDocumentCleaner, added remove_empty_rows & remove_empty_columns to optionally remove rows and columns. Also added keep_id to optionally allow for keeping the original document ID.OpenAPIServiceConnector to support and be compatible with the new ChatMessage format.ExtractedTableAnswer dataclass and the dataframe field in the Document dataclass are deprecated and will be removed in Haystack 2.11.0. Check out the GitHub discussion for motivation and details.DOCXToDocument component now skips comment blocks in DOCX files that previously caused errors.JSONConverter to properly skip converting JSON files that are not utf-8 encoded.PDFMinerToDocument convert function to to double new lines between container_text so that passages can later by DocumentSplitter.OpenAIChatGenerator streaming response tool call processing: The logic now scans all chunks to correctly identify the first chunk with tool calls, ensuring accurate payload construction and preventing errors when tool call data isn't confined to the initial chunk.Nothing published for this version
Nothing published for this version
The refactoring of the ChatMessage data class includes some breaking changes involving ChatMessage creation and accessing attributes. If you have a Pi…
We are introducing the Tool, a simple and unified abstraction for representing tools in Haystack, and the ToolInvoker, which executes tool calls prepared by LLMs. These features make it easy to integrate tool calling into your Haystack pipelines, enabling seamless interaction with tools when used with components like OpenAIChatGenerator and HuggingFaceAPIChatGenerator. Here's how you can use them:
def dummy_weather_function(city: str):
return f"The weather in {city} is 20 degrees."
tool = Tool(
name="weather_tool",
description="A tool to get the weather",
function=dummy_weather_function,
parameters={
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
}
)
pipeline = Pipeline()
pipeline.add_component("llm", OpenAIChatGenerator(model="gpt-4o-mini", tools=[tool]))
pipeline.add_component("tool_invoker", ToolInvoker(tools=[tool]))
pipeline.connect("llm.replies", "tool_invoker.messages")
message = ChatMessage.from_user("How is the weather in Berlin today?")
result = pipeline.run({"llm": {"messages": [message]}})
Use Components as Tools
As an abstraction of Tool, ComponentTool allows LLMs to interact directly with components like web search, document processing, or custom user components. It simplifies schema generation and type conversion, making it easy to expose complex component functionality to LLMs.
# Create a tool from the component
tool = ComponentTool(
component=SerperDevWebSearch(api_key=Secret.from_env_var("SERPERDEV_API_KEY"), top_k=3),
name="web_search", # Optional: defaults to "serper_dev_web_search"
description="Search the web for current information on any topic" # Optional: defaults to component docstring
)
RecursiveDocumentSplitterRecursiveDocumentSplitter introduces a smarter way to split text. It uses a set of separators to divide text recursively, starting with the first separator. If chunks are still larger than the specified size, the splitter moves to the next separator in the list. This approach ensures efficient and granular text splitting for improved processing.
from haystack.components.preprocessors import RecursiveDocumentSplitter
splitter = RecursiveDocumentSplitter(split_length=260, split_overlap=0, separators=["\n\n", "\n", ".", " "])
doc_chunks = splitter.run([Document(content="...")])
ChatMessage dataclassChatMessage dataclass has been refactored to improve flexibility and compatibility. As part of this update, the content attribute has been removed and replaced with a new text property for accessing the ChatMessage's textual value. This change ensures future-proofing and better support for features like tool calls and their results. For details on the new API and migration steps, see the ChatMessage documentation. If you have any questions about this refactoring, feel free to let us know in this Github discussion.
ChatMessage data class includes some breaking changes involving ChatMessage creation and accessing attributes. If you have a Pipeline containing a ChatPromptBuilder, serialized with haystack-ai =< 2.9.0, deserialization may break. For detailed information about the changes and how to migrate, see the ChatMessage documentation.converter init argument from PyPDFToDocument. Use other init arguments instead, or create a custom component.SentenceWindowRetriever output key context_documents now outputs a List[Document] containing the retrieved documents and the context windows ordered by split_idx_start.store_full_path to False in convertersIntroduced the ComponentTool, a new tool that wraps Haystack components, allowing them to be utilized as tools for LLMs (various ChatGenerators). This ComponentTool supports automatic tool schema generation, input type conversion, and offers support for components with run methods that have input types:
List[str])List[Document])List[Document], str etc.)Example usage:
from haystack import component, Pipeline
from haystack.tools import ComponentTool
from haystack.components.websearch import SerperDevWebSearch
from haystack.utils import Secret
from haystack.components.tools.tool_invoker import ToolInvoker
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
# Create a SerperDev search component
search = SerperDevWebSearch(api_key=Secret.from_env_var("SERPERDEV_API_KEY"), top_k=3)
# Create a tool from the component
tool = ComponentTool(
component=search,
name="web_search", # Optional: defaults to "serper_dev_web_search"
description="Search the web for current information on any topic" # Optional: defaults to component docstring
)
# Create pipeline with OpenAIChatGenerator and ToolInvoker
pipeline = Pipeline()
pipeline.add_component("llm", OpenAIChatGenerator(model="gpt-4o-mini", tools=[tool]))
pipeline.add_component("tool_invoker", ToolInvoker(tools=[tool]))
# Connect components
pipeline.connect("llm.replies", "tool_invoker.messages")
message = ChatMessage.from_user("Use the web search tool to find information about Nikola Tesla")
# Run pipeline
result = pipeline.run({"llm": {"messages": [message]}})
print(result)
Add XLSXToDocument converter that loads an Excel file using Pandas + openpyxl and by default converts each sheet into a separate Document in CSV format.
Added a new store_full_path parameter to the __init__ methods of PyPDFToDocument and AzureOCRDocumentConverter. The default value is True, which stores the full file path in the metadata of the output documents. When set to False, only the file name is stored.
Add a new experimental component ToolInvoker. This component invokes tools based on tool calls prepared by Language Models and returns the results as a list of ChatMessage objects with tool role.
Adding a RecursiveSplitter, which uses a set of separators to split text recursively. It attempts to divide the text using the first separator, and if the resulting chunks are still larger than the specified size, it moves to the next separator in the list.
Added a create_tool_from_function function to create a Too instance from a function, with automatic generation of name, description and parameters. Added a tool decorator to achieve the same result.
Add support for Tools in the Hugging Face API Chat Generator.
Changed the ChatMessage dataclass to support different types of content, including tool calls, and tool call results.
Add support for Tools in the OpenAI Chat Generator.
Added a new Tool dataclass to represent a tool for which Language Models can prepare calls.
Added the component StringJoiner to join strings from different components to a list of strings.
Added default_headers parameter to AzureOpenAIDocumentEmbedder and AzureOpenAITextEmbedder.
Add token argument to NamedEntityExtractor to allow usage of private Hugging Face models.
Add the from_openai_dict_format class method to the ChatMessage class. It allows you to create a ChatMessage from a dictionary in the format that OpenAI's Chat API expects.
Add a testing job to check that all packages can be imported successfully. This should help detect several issues, such as forgetting to use a forward reference for a type hint coming from a lazy import.
DocumentJoiner methods _concatenate() and _distribution_based_rank_fusion() were converted to static methods.
Improve serialization and deserialization of callables. We now allow serialization of class methods and static methods and explicitly prohibit serialization of instance methods, lambdas, and nested functions.
Added new initialization parameters to the PyPDFToDocument component to customize the text extraction process from PDF files.
Reorganized the document store test suite to isolate dataframe filter tests. This change prepares for potential future deprecation of the Document class's dataframe field.
Move Tool to a new dedicated tools package. Refactor Tool serialization and deserialization to make it more flexible and include type information.
The NLTKDocumentSplitter was merged into the DocumentSplitter which now provides the same functionality as the NLTKDocumentSplitter. The split_by="sentence" now uses a custom sentence boundary detection based on the nltk library. The previous sentence behaviour can still be achieved by split_by="period".
Improved deserialization of callables by using importlib instead of sys.modules. This change allows importing local functions and classes that are not in sys.modules when deserializing callable.
Change OpenAIDocumentEmbedder to keep running if a batch fails embedding. Now OpenAI returns an error we log that error and keep processing following batches.
The NLTKDocumentSplitter will be deprecated and will be removed in the next release. The DocumentSplitter will instead support the functionality of the NLTKDocumentSplitter.
The function role and ChatMessage.from_function class method have been deprecated and will be removed in Haystack 2.10.0. ChatMessage.from_function also attempts to produce a valid tool message. For more information, see the documentation: https://docs.haystack.deepset.ai/docs/chatmessage
The SentenceWindowRetriever output of context_documents changed. Instead of a List[List[Document], the output is a List[Document], where the documents are ordered by split_idx_start value.
Add missing stream mime type assignment to the LinkContentFetcher for the single url scenario.
Previously, the pipelines that use FileTypeRouter could fail if they received a single URL as an input.
OpenAIChatGenerator no longer passes tools to the OpenAI client if none are provided. Previously, a null value was passed. This change improves compatibility with OpenAI-compatible APIs that do not support tools.
ByteStream now truncates the data to 100 bytes in the string representation to avoid excessive log output.
Make the HuggingFaceLocalChatGenerator compatible with the new ChatMessage format, by converting the messages to the format expected by HuggingFace.
Serialize the chat_template parameter.
Moved the NLTK download of DocumentSplitter and NLTKDocumentSplitter to warm_up(). This prevents calling to an external API during instantiation. If a DocumentSplitter or NLTKDocumentSplitter is used for sentence splitting outside of a pipeline, warm_up() now needs to be called before running the component.
PDFMinerToDocument now creates documents with id based on converted text and metadata. Before, PDFMinerToDocument did not consider the document's meta field when generating the document's id.
Pin OpenAI client to >=1.56.1 to avoid issues related to changes in the httpx library.
PyPDFToDocument now creates documents with id based on converted text and metadata. Before it didn't take the meta data into account.
Fixes issues with deserialization of components in multi-threaded environments.
Nothing published for this version
Pin OpenAI client to \>=1.56.1 to avoid issues related to changes in the httpx library.
PyPDFToDocument now creates documents with id based on converted text and meta data. Before it didn't take the meta data into account.
Fixes issues with deserialization of components in multi-threaded environments.
Pin OpenAI client to \>=1.56.1 to avoid issues related to changes in the httpx library.
Remove is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.
is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.DALLEImageGenerator component, bringing image generation with OpenAI's DALL-E to the Haystack
from haystack.components.generators import DALLEImageGenerator
image_generator = DALLEImageGenerator()
response = image_generator.run("Show me a picture of a black cat.")
print(response)
PDFMinerToDocument and PyPDFToDocument to indicate when a processed PDF file has no content. This can happen if the PDF file is a scanned image. Also added an explicit check and warning message to the DocumentSplitter that warns the user that empty Documents are skipped. This behavior was already occurring, but now its clearer through logs that this is happening.MetaFieldGroupingRanker component that reorders documents by grouping them based on metadata keys. This can be useful for pre-processing Documents before feeding them to an LLM.store_full_path parameter to the __init__ methods of the following converters:
JSONConverter, CSVToDocument, DOCXToDocument, HTMLToDocument MarkdownToDocument, PDFMinerToDocument, PPTXToDocument, TikaDocumentConverter, PyPDFToDocument , AzureOCRDocumentConverter and TextFileToDocument. The default value is True, which stores full file path in the metadata of the output documents. When set to False, only the file name is stored.OpenAPI, allow both switching SSL verification off and specifying a certificate authority to use for it.PromptBuilder and ChatPromptBuilder. By passing required_variables="*" you can automatically set all variables in the prompt to be required.ChatMessage data class constructor with specific class methods (ChatMessage.from_user, ChatMessage.from_assistant, etc.).SentenceTransformersDiversityRanker. MMR scores are calculated for each document based on their relevance to the query and diversity from already selected documents.ConditionalRouter component, enabling default/fallback routing behavior when certain inputs are not provided at runtime. This enhancement allows for more flexible pipeline configurations with graceful handling of missing parameters.DocumentSplitter, which will split the document at n.OpenAIDocumentEmbedder to keep running if a batch fails embedding. Now OpenAI returns an error we log that error and keep processing following batches.PyPDFToDocument component to customize the text extraction process from PDF files.ChatMessage.content with ChatMessage.text across the codebase. This is done in preparation for the removal of content in Haystack 2.9.0.store_full_path parameter in converters will change to False in Haysatck 2.9.0 to enhance privacy.ChatMessage data class will be refactored to make it more flexible and future-proof. As part of this change, the <span class="title-ref">content</span> attribute will be removed. A new text property has been introduced to provide access to the textual value of the ChatMessage. To ensure a smooth transition, start using the text property now in place of content.converter parameter in the PyPDFToDocument component is deprecated and will be removed in Haystack 2.9.0. For in-depth customization of the conversion process, consider implementing a custom component. Additional high-level customization options will be added in the future.context_documents in SentenceWindowRetriever will change in the next release. Instead of a List[List[Document]], the output will be a List[Document], where the documents are ordered by split_idx_start.Fix DocumentCleaner not preserving all Document fields when run
Fix DocumentJoiner failing when ran with an empty list of Documents
For the NLTKDocumentSplitter we are updating how chunks are made when splitting by word and sentence boundary is respected. Namely, to avoid fully subsuming the previous chunk into the next one, we ignore the first sentence from that chunk when calculating sentence overlap. i.e. we want to avoid cases of Doc1 = [s1, s2], Doc2 = [s1, s2, s3].
Finished adding function support for this component by updating the _split_into_units function and added the splitting_function init parameter.
Add specific to_dict method to overwrite the underlying one from DocumentSplitter. This is needed to properly save the settings of the component to yaml.
Fix OpenAIChatGenerator and OpenAIGenerator crashing when using a <span class="title-ref">streaming_callback</span> and generation_kwargs contain {"stream_options": {"include_usage": True}}.
Fix tracing Pipeline with cycles to correctly track components execution
When meta is passed into AnswerBuilder.run(), it is now merged into GeneratedAnswer meta
Fix DocumentSplitter to handle custom splitting_function without requiring split_length. Previously the splitting_function provided would not override other settings.
Remove is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.
is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.DALLEImageGenerator component, bringing image generation with OpenAI's DALL-E to the Haystack
`python from haystack.components.generators import DALLEImageGenerator image_generator = DALLEImageGenerator() response = image_generator.run("Show me a picture of a black cat.") print(response)`PDFMinerToDocument and PyPDFToDocument to indicate when a processed PDF file has no content. This can happen if the PDF file is a scanned image. Also added an explicit check and warning message to the DocumentSplitter that warns the user that empty Documents are skipped. This behavior was already occurring, but now its clearer through logs that this is happening.MetaFieldGroupingRanker component that reorders documents by grouping them based on metadata keys. This can be useful for pre-processing Documents before feeding them to an LLM.store_full_path parameter to the __init__ methods of the following converters:
JSONConverter, CSVToDocument, DOCXToDocument, HTMLToDocument MarkdownToDocument, PDFMinerToDocument, PPTXToDocument, TikaDocumentConverter, PyPDFToDocument , AzureOCRDocumentConverter and TextFileToDocument. The default value is True, which stores full file path in the metadata of the output documents. When set to False, only the file name is stored.OpenAPI, allow both switching SSL verification off and specifying a certificate authority to use for it.PromptBuilder and ChatPromptBuilder. By passing required_variables="*" you can automatically set all variables in the prompt to be required.ChatMessage data class constructor with specific class methods (ChatMessage.from_user, ChatMessage.from_assistant, etc.).SentenceTransformersDiversityRanker. MMR scores are calculated for each document based on their relevance to the query and diversity from already selected documents.ConditionalRouter component, enabling default/fallback routing behavior when certain inputs are not provided at runtime. This enhancement allows for more flexible pipeline configurations with graceful handling of missing parameters.DocumentSplitter, which will split the document at n.OpenAIDocumentEmbedder to keep running if a batch fails embedding. Now OpenAI returns an error we log that error and keep processing following batches.PyPDFToDocument component to customize the text extraction process from PDF files.ChatMessage.content with ChatMessage.text across the codebase. This is done in preparation for the removal of content in Haystack 2.9.0.store_full_path parameter in converters will change to False in Haysatck 2.9.0 to enhance privacy.ChatMessage data class will be refactored to make it more flexible and future-proof. As part of this change, the <span class="title-ref">content</span> attribute will be removed. A new text property has been introduced to provide access to the textual value of the ChatMessage. To ensure a smooth transition, start using the text property now in place of content.converter parameter in the PyPDFToDocument component is deprecated and will be removed in Haystack 2.9.0. For in-depth customization of the conversion process, consider implementing a custom component. Additional high-level customization options will be added in the future.context_documents will change in the next release. Instead of a List[List[Document]], the output will be a List[Document], where the documents are ordered by split_idx_start.Fix DocumentCleaner not preserving all Document fields when run
Fix DocumentJoiner failing when ran with an empty list of Documents
For the NLTKDocumentSplitter we are updating how chunks are made when splitting by word and sentence boundary is respected. Namely, to avoid fully subsuming the previous chunk into the next one, we ignore the first sentence from that chunk when calculating sentence overlap. i.e. we want to avoid cases of Doc1 = [s1, s2], Doc2 = [s1, s2, s3].
Finished adding function support for this component by updating the _split_into_units function and added the splitting_function init parameter.
Add specific to_dict method to overwrite the underlying one from DocumentSplitter. This is needed to properly save the settings of the component to yaml.
Fix OpenAIChatGenerator and OpenAIGenerator crashing when using a <span class="title-ref">streaming_callback</span> and generation_kwargs contain {"stream_options": {"include_usage": True}}.
Fix tracing Pipeline with cycles to correctly track components execution
When meta is passed into AnswerBuilder.run(), it is now merged into GeneratedAnswer meta
Fix DocumentSplitter to handle custom splitting_function without requiring split_length. Previously the splitting_function provided would not override other settings.
Remove is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.
is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.DALLEImageGenerator component, bringing image generation with OpenAI's DALL-E to the Haystack
`python from haystack.components.generators import DALLEImageGenerator image_generator = DALLEImageGenerator() response = image_generator.run("Show me a picture of a black cat.") print(response)`PDFMinerToDocument and PyPDFToDocument to indicate when a processed PDF file has no content. This can happen if the PDF file is a scanned image. Also added an explicit check and warning message to the DocumentSplitter that warns the user that empty Documents are skipped. This behavior was already occurring, but now its clearer through logs that this is happening.MetaFieldGroupingRanker component that reorders documents by grouping them based on metadata keys. This can be useful for pre-processing Documents before feeding them to an LLM.store_full_path parameter to the __init__ methods of the following converters:
JSONConverter, CSVToDocument, DOCXToDocument, HTMLToDocument MarkdownToDocument, PDFMinerToDocument, PPTXToDocument, TikaDocumentConverter, PyPDFToDocument , AzureOCRDocumentConverter and TextFileToDocument. The default value is True, which stores full file path in the metadata of the output documents. When set to False, only the file name is stored.OpenAPI, allow both switching SSL verification off and specifying a certificate authority to use for it.PromptBuilder and ChatPromptBuilder. By passing required_variables="*" you can automatically set all variables in the prompt to be required.ChatMessage data class constructor with specific class methods (ChatMessage.from_user, ChatMessage.from_assistant, etc.).SentenceTransformersDiversityRanker. MMR scores are calculated for each document based on their relevance to the query and diversity from already selected documents.ConditionalRouter component, enabling default/fallback routing behavior when certain inputs are not provided at runtime. This enhancement allows for more flexible pipeline configurations with graceful handling of missing parameters.DocumentSplitter, which will split the document at nOpenAIDocumentEmbedder to keep running if a batch fails embedding. Now OpenAI returns an error we log that error and keep processing following batches.PyPDFToDocument component to customize the text extraction process from PDF files.ChatMessage.content with ChatMessage.text across the codebase. This is done in preparation for the removal of content in Haystack 2.9.0.store_full_path parameter will change to <span class="title-ref">False</span> in Haysatck 2.9.0 to enhance privacy.store_full_path parameter in converters will change to False in Haysatck 2.9.0 to enhance privacy.ChatMessage data class will be refactored to make it more flexible and future-proof. As part of this change, the <span class="title-ref">content</span> attribute will be removed. A new text property has been introduced to provide access to the textual value of the ChatMessage. To ensure a smooth transition, start using the text property now in place of content.converter parameter in the PyPDFToDocument component is deprecated and will be removed in Haystack 2.9.0. For in-depth customization of the conversion process, consider implementing a custom component. Additional high-level customization options will be added in the future.Fix DocumentCleaner not preserving all Document fields when run
Fix DocumentJoiner failing when ran with an empty list of Documents
For the NLTKDocumentSplitter we are updating how chunks are made when splitting by word and sentence boundary is respected. Namely, to avoid fully subsuming the previous chunk into the next one, we ignore the first sentence from that chunk when calculating sentence overlap. i.e. we want to avoid cases of Doc1 = [s1, s2], Doc2 = [s1, s2, s3].
Finished adding function support for this component by updating the _split_into_units function and added the splitting_function init parameter.
Add specific to_dict method to overwrite the underlying one from DocumentSplitter. This is needed to properly save the settings of the component to yaml.
Fix OpenAIChatGenerator and OpenAIGenerator crashing when using a <span class="title-ref">streaming_callback</span> and generation_kwargs contain {"stream_options": {"include_usage": True}}.
Fix tracing Pipeline with cycles to correctly track components execution
When meta is passed into AnswerBuilder.run(), it is now merged into GeneratedAnswer meta
Fix DocumentSplitter to handle custom splitting_function without requiring split_length. Previously the splitting_function provided would not override other settings.
Remove is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.
is_greedy deprecated argument from @component decorator. Change the Variadic input of your Component to GreedyVariadic instead.DALLEImageGenerator component, bringing image generation with OpenAI's DALL-E to the Haystack
`python from haystack.components.generators import DALLEImageGenerator image_generator = DALLEImageGenerator() response = image_generator.run("Show me a picture of a black cat.") print(response)`PDFMinerToDocument and PyPDFToDocument to indicate when a processed PDF file has no content. This can happen if the PDF file is a scanned image. Also added an explicit check and warning message to the DocumentSplitter that warns the user that empty Documents are skipped. This behavior was already occurring, but now its clearer through logs that this is happening.MetaFieldGroupingRanker component that reorders documents by grouping them based on metadata keys. This can be useful for pre-processing Documents before feeding them to an LLM.store_full_path parameter to the __init__ methods of the following converters:
JSONConverter, CSVToDocument, DOCXToDocument, HTMLToDocument MarkdownToDocument, PDFMinerToDocument, PPTXToDocument, TikaDocumentConverter and TextFileToDocument. The default value is True, which stores full file path in the metadata of the output documents. When set to False, only the file name is stored.OpenAPI, allow both switching SSL verification off and specifying a certificate authority to use for it.PromptBuilder and ChatPromptBuilder. By passing required_variables="*" you can automatically set all variables in the prompt to be required.SentenceTransformersDiversityRanker. MMR scores are calculated for each document based on their relevance to the query and diversity from already selected documents.ConditionalRouter component, enabling default/fallback routing behavior when certain inputs are not provided at runtime. This enhancement allows for more flexible pipeline configurations with graceful handling of missing parameters.DocumentSplitter, which will split the document at nOpenAIDocumentEmbedder to keep running if a batch fails embedding. Now OpenAI returns an error we log that error and keep processing following batches.store_full_path parameter will change to <span class="title-ref">False</span> in Haysatck 2.9.0 to enhance privacy.Fix DocumentCleaner not preserving all Document fields when run
Fix DocumentJoiner failing when ran with an empty list of Documents
For the NLTKDocumentSplitter we are updating how chunks are made when splitting by word and sentence boundary is respected. Namely, to avoid fully subsuming the previous chunk into the next one, we ignore the first sentence from that chunk when calculating sentence overlap. i.e. we want to avoid cases of Doc1 = [s1, s2], Doc2 = [s1, s2, s3].
Finished adding function support for this component by updating the _split_into_units function and added the splitting_function init parameter.
Add specific to_dict method to overwrite the underlying one from DocumentSplitter. This is needed to properly save the settings of the component to yaml.
Fix OpenAIChatGenerator and OpenAIGenerator crashing when using a <span class="title-ref">streaming_callback</span> and generation_kwargs contain {"stream_options": {"include_usage": True}}.
Fix tracing Pipeline with cycles to correctly track components execution
When meta is passed into AnswerBuilder.run(), it is now merged into GeneratedAnswer meta
Fix DocumentSplitter to handle custom splitting_function without requiring split_length. Previously the splitting_function provided would not override other settings.
The deprecated default converter class haystack.components.converters.pypdf.DefaultConverter used by PyPDFToDocument has been removed.
Pipeline.run() logic to better handle cyclesPipeline.run() internal logic has been heavily reworked to be more robust and reliable than before. This new implementation makes it easier to run Pipelines that have cycles in their graph. It also fixes some corner cases in Pipelines that don't have any cycle.
LoggingTracerWith the new LoggingTracer, users can inspect the logs in real-time to see everything that is happening in their Pipelines. This feature aims to improve the user experience during experimentation and prototyping.
import logging
from haystack import tracing
from haystack.tracing.logging_tracer import LoggingTracer
logging.basicConfig(format="%(levelname)s - %(name)s - %(message)s", level=logging.WARNING)
logging.getLogger("haystack").setLevel(logging.DEBUG)
tracing.tracer.is_content_tracing_enabled = True # to enable tracing/logging content (inputs/outputs)
tracing.enable_tracing(LoggingTracer())
Removed Pipeline init argument debug_path. We do not support this anymore.
Removed Pipeline init argument max_loops_allowed. Use max_runs_per_component instead.
Removed PipelineMaxLoops exception. Use PipelineMaxComponentRuns instead.
The deprecated default converter class haystack.components.converters.pypdf.DefaultConverter used by PyPDFToDocument has been removed.
Pipeline YAMLs from haystack<2.7.0 that use the default converter must be updated in the following manner:
# Old
components:
Comp1:
init_parameters:
converter:
type: haystack.components.converters.pypdf.DefaultConverter
type: haystack.components.converters.pypdf.PyPDFToDocument
# New
components:
Comp1:
init_parameters:
converter: null
type: haystack.components.converters.pdf.PDFToTextConverter
Pipeline YAMLs from haystack<2.7.0 that use custom converter classes can be upgraded by simply loading them with haystack==2.6.x and saving them to YAML again.
Pipeline.connect() will now raise a PipelineConnectError if sender and receiver are the same Component. We do not support this use case anymore.
Added component StringJoiner to join strings from different components to a list of strings.
Improved serialization/deserialization errors to provide extra context about the delinquent components when possible.
Enhanced DOCX converter to support table extraction in addition to paragraph content. The converter supports both CSV and Markdown table formats, providing flexible options for representing tabular data extracted from DOCX documents.
Added a new parameter additional_mimetypes to the FileTypeRouter component. This allows users to specify additional MIME type mappings, ensuring correct file classification across different runtime environments and Python versions.
Introduce a LoggingTracer, that sends all traces to the logs.
It can enabled as follows:
import logging
from haystack import tracing
from haystack.tracing.logging_tracer import LoggingTracer
logging.basicConfig(format="%(levelname)s - %(name)s - %(message)s", level=logging.WARNING)
logging.getLogger("haystack").setLevel(logging.DEBUG)
tracing.tracer.is_content_tracing_enabled = True # to enable tracing/logging content (inputs/outputs)
tracing.enable_tracing(LoggingTracer())
Fundamentally rework the internal logic of Pipeline.run(). The rework makes it more reliable and covers more use cases. We fixed some issues that made Pipelines with cycles unpredictable and with unclear Components execution order.
Each tracing span of a component run is now attached with the pipeline run span object. This allows users to trace the execution of multiple pipeline runs concurrently.
streaming_callback run parameter to HuggingFaceAPIGenerator and HuggingFaceLocalGenerator to allow users to pass a callback function that will be called after each chunk of the response is generated.SentenceWindowRetriever now supports the window_size parameter at run time, overwriting the value set in the constructor.ConditionalRouter. Setting validate_output_type to True will enable a check to verify if the actual output of a route returns the declared type. If it doesn't match a ValueError is raised.numpy usage to speed up imports.FileTypeRouter, particularly for Microsoft Office file formats like .docx and .pptx. This enhancement ensures more consistent behavior across different environments, including AWS Lambda functions and systems without pre-installed office suites.FiletypeRouter now supports passing metadata (meta) in the run method. When metadata is provided, the sources are internally converted to ByteStream objects and the metadata is added. This new parameter simplifies working with preprocessing/indexing pipelines.SentenceTransformersDocumentEmbedder now supports config_kwargs for additional parameters when loading the model configurationSentenceTransformersTextEmbedder now supports config_kwargs for additional parameters when loading the model configurationnumpy was pinned to <2.0 to avoid compatibility issues in several core integrations. This pin has been removed, and haystack can work with both numpy 1.x and 2.x. If necessary, we will pin numpy version in specific core integrations that require it.DefaultConverter class used by the PyPDFToDocument component has been deprecated. Its functionality will be merged into the component in 2.7.0.str, int, float, bool, list, dict, set, tuple or None.HuggingFaceAPIGenerator component to prevent import errors.PyPDFToDocument component to prevent the default converter from being serialized unnecessarily.PyPDFConverter that broke the deserialization of pre 2.6.0 YAMLs.The deprecated default converter class haystack.components.converters.pypdf.DefaultConverter used by PyPDFToDocument has been removed.
Pipeline.run() logic to better handle cyclesPipeline.run() internal logic has been heavily reworked to be more robust and reliable than before. This new implementation makes it easier to run Pipelines that have cycles in their graph. It also fixes some corner cases in Pipelines that don't have any cycle.
LoggingTracerWith the new LoggingTracer, users can inspect in the logs everything that is happening in their Pipelines in real time. This feature aims to improve the user experience during experimentation and prototyping.
Removed Pipeline init argument debug_path. We do not support this anymore.
Removed Pipeline init argument max_loops_allowed. Use max_runs_per_component instead.
Removed PipelineMaxLoops exception. Use PipelineMaxComponentRuns instead.
The deprecated default converter class haystack.components.converters.pypdf.DefaultConverter used by PyPDFToDocument has been removed.
Pipeline YAMLs from haystack<2.7.0 that use the default converter must be updated in the following manner:
# Old
components:
Comp1:
init_parameters:
converter:
type: haystack.components.converters.pypdf.DefaultConverter
type: haystack.components.converters.pypdf.PyPDFToDocument
# New
components:
Comp1:
init_parameters:
converter: null
type: haystack.components.converters.pdf.PDFToTextConverter
Pipeline YAMLs from haystack<2.7.0 that use custom converter classes can be upgraded by simply loading them with haystack==2.6.x and saving them to YAML again.
Pipeline.connect() will now raise a PipelineConnectError if sender and receiver are the same Component. We do not support this use case anymore.
Added component StringJoiner to join strings from different components to a list of strings.
Improved serialization/deserialization errors to provide extra context about the delinquent components when possible.
Enhanced DOCX converter to support table extraction in addition to paragraph content. The converter supports both CSV and Markdown table formats, providing flexible options for representing tabular data extracted from DOCX documents.
Added a new parameter additional_mimetypes to the FileTypeRouter component.
This allows users to specify additional MIME type mappings, ensuring correct
file classification across different runtime environments and Python versions.
Introduce a LoggingTracer, that sends all traces to the logs.
It can enabled as follows:
import logging
from haystack import tracing
from haystack.tracing.logging_tracer import LoggingTracer
logging.basicConfig(format="%(levelname)s - %(name)s - %(message)s", level=logging.WARNING)
logging.getLogger("haystack").setLevel(logging.DEBUG)
tracing.tracer.is_content_tracing_enabled = True # to enable tracing/logging content (inputs/outputs)
tracing.enable_tracing(LoggingTracer())
Fundamentally rework the internal logic of Pipeline.run(). The rework makes it more reliable and covers more use cases. We fixed some issues that made Pipelines with cycles unpredictable and with unclear Components execution order.
Each tracing span of a component run is now attached with the pipeline run span object. This allows users to trace the execution of multiple pipeline runs concurrently.
streaming_callback run parameter to HuggingFaceAPIGenerator and HuggingFaceLocalGenerator to allow users to pass a callback function that will be called after each chunk of the response is generated.SentenceWindowRetriever now supports the window_size parameter at run time, overwriting the value set in the constructor.ConditionalRouter. Setting validate_output_type to True will enable a check to verify if the actual output of a route returns the declared type. If it doesn't match a ValueError is raised.numpy usage to speed up imports.FileTypeRouter, particularly for Microsoft Office file formats like .docx and .pptx. This enhancement ensures more consistent behavior across different environments, including AWS Lambda functions and systems without pre-installed office suites.FiletypeRouter now supports passing metadata (meta) in the run method. When metadata is provided, the sources are internally converted to ByteStream objects and the metadata is added. This new parameter simplifies working with preprocessing/indexing pipelines.SentenceTransformersDocumentEmbedder now supports config_kwargs for additional parameters when loading the model configurationSentenceTransformersTextEmbedder now supports config_kwargs for additional parameters when loading the model configurationnumpy was pinned to <2.0 to avoid compatibility issues in several core integrations. This pin has been removed, and haystack can work with both numpy 1.x and 2.x. If necessary, we will pin numpy version in specific core integrations that require it.DefaultConverter class used by the PyPDFToDocument component has been deprecated. Its functionality will be merged into the component in 2.7.0.str, int, float, bool, list, dict, set, tuple or None.HuggingFaceAPIGenerator component to prevent import errors.PyPDFToDocument component to prevent the default converter from being serialized unnecessarily.PyPDFConverter that broke the deserialization of pre 2.6.0 YAMLs.Revert change to PyPDFConverter that broke the deserialization of pre 2.6.0 YAMLs.
Your coding agent can read these notes before it upgrades. Set up the MCP server →