NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #5108 most downloaded on PyPI
LLM framework to build customizable, production-ready LLM applications. Connect components (models, vector DBs, file converters) to pipelines or agents that can interact with your data.
Last release today
05 Oct 2026
Ships fairly regularly
a new release about every 8 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
231 releases · first in 2023
Nothing published for this version
Nothing published for this version
Nothing published for this version
One column per month.
Nothing published for this version
Nothing published for this version
anyio is installed transitively through httpx and openai ; versions before 4.14.2 are affected by CVE-2026-63374 ( GHSA-82r6-8w77-94w6 ).
SentenceWindowRetriever now queries the Document Store once per run or run_async call instead of once per retrieved document. Fewer calls mean lower latency and less load on the Document Store and outputs have not changed, so existing pipelines get faster without code changes.
InMemoryDocumentStore.bm25_retrieval and InMemoryBM25Retriever with BM25L (the default) or BM25Plus now return only documents that contain at least one query term, with or without scale_score. Previously, documents without any query term were returned with a positive score and filled up top_k. As a result, retrieval can now return fewer than top_k documents, or none at all. The delta lower bound now applies only to query terms that occur in the document, as in the original BM25L/BM25+ definitions. So scores of matching documents are lower than before whenever the query has terms that a document does not contain. If you filter results with a fixed score threshold, re-check the threshold. BM25Okapi is not affected.TextCleaner.run now raises a clear TypeError when texts is not a list or any element is not a str, instead of failing later or producing unexpected results.anyio>=4.14.2. anyio is installed transitively through httpx and openai; versions before 4.14.2 are affected by CVE-2026-63374 (GHSA-82r6-8w77-94w6).Fixed BM25 tokenization in InMemoryDocumentStore for Chinese, Japanese and Korean text. The default bm25_tokenization_regex now splits CJK characters into one token each, so bare-term queries can match words inside longer unspaced runs. Text is also NFC-normalized before tokenization, so composed and decomposed spellings produce the same tokens.
Fix Pipeline.connect() raising TypeError: object of type 'ellipsis' has no len() when one of the sockets is annotated with Callable[..., T]. A Callable[..., T] now matches callables with any parameters, in both directions, as long as the return types are compatible.
Fixed ChatPromptBuilder dropping content parts when the template is a list of ChatMessage objects. Only the first TextContent part was rendered, and every other part - additional texts, images, files -was removed from the rendered prompt without a warning. Template variables used in those dropped parts were not detected either, so they were not exposed as inputs. Now all TextContent parts are rendered and the remaining content parts are passed through unchanged.
Fixed CompactionHook.close() and close_async() to release the token counter's resources as well as the compactor's. Previously, resources such as OpenAITokenCounter's HTTP client remained open after the hook or Agent was closed.
ConfirmationHook no longer drops the messages that come after the last user or tool message when it rewrites the conversation, such as a system message added by another hook or an assistant answer followed by an on_exit reminder. Before, those messages were removed from the history the Agent sends to the model.
DocumentToImageContent no longer raises a KeyError for the whole batch when a document points to a PDF page that cannot be converted, because the page is out of range or the PDF cannot be read. That document now gets None in image_contents, like other invalid documents, and the rest of the batch is still converted. The logged warnings now include the path of the PDF file.
Fixed DOCXToDocument breaking a Markdown table when a cell contains a pipe or spans several paragraphs. A pipe was emitted as a column separator, and a cell's line break ended the row in the middle of it. Pipes are now escaped and line breaks inside a cell are collapsed to a space. The csv table format is unchanged: it already kept such cells intact by quoting them.
Fixed DOCXToDocument dropping hyperlink addresses inside tables. With link_format set to markdown or plain, links in table cells were written as their display text only, while links in body paragraphs kept their address. Links in table cells are now formatted the same way, in both the markdown and csv table formats. The default link_format="none" output is unchanged.
Fixed ConditionalRouter.from_dict mutating the caller's routes data in place: serialized output_type strings were deserialized directly inside the caller's dictionaries, so reusing the same serialized pipeline dict afterwards yielded already-deserialized type objects. from_dict now works on a copy and the caller's data is left untouched. BranchJoiner.from_dict received the same treatment.
LinkContentFetcher no longer adds its timeout and follow_redirects defaults to the client_kwargs dictionary passed by the caller. The dictionary is now copied before the defaults are applied, so reusing one HTTP client configuration across components no longer leaks these defaults.
Fix MetaFieldRanker to return no documents when missing_meta="drop" and all documents lack the ranking field or have a None value.
MetaFieldRanker now treats a Document whose meta_field value is None the same as a Document that is missing the field, applying the missing_meta setting to it. Previously a single None value made sorting fail, so the ranker logged a warning and returned all Documents in their original order, ignoring missing_meta="drop" as well.
MockTextEmbedder and MockDocumentEmbedder now accept non-positive dimension values when embedding or embedding_fn is provided, matching the documented behavior. The default deterministic embedding mode still requires a positive dimension.
ChatMessage.from_openai_dict_format now accepts tool-call arguments that are already a dictionary instead of raising a TypeError. Some OpenAI-compatible servers send a parsed object rather than a JSON string.
MetaFieldGroupingRanker now treats a group_by or subgroup_by value of None as missing, the same way it already treats sort_docs_by. Documents with None go to the end with the other ungrouped documents, instead of forming a group named "None" that also absorbed documents whose value was the string "None".
Fixed JSONConverter failing with ValueError when content_key contains numeric or boolean scalar values. These values are now converted to strings before creating the Document, while null values remain unchanged.
Fixed SentenceSplitter losing the whitespace between two sentences when the first one ends with a closing quote, as in He said "Hi." Bye.. Those characters were missing from the chunk text and shifted the split_idx_start offset of every following chunk, so chunks could no longer be mapped back onto the original text. This affects all components that split on sentences, such as DocumentSplitter, RecursiveDocumentSplitter, MarkdownHeaderSplitter and EmbeddingBasedDocumentSplitter. Chunk boundaries change for text that contains quoted sentences, so re-indexing an existing corpus produces different chunks than before.
Fixed LLMMetadataExtractor incorrectly treating a chat generator output that carries its own error field as a failed LLM call, which sent every document to failed_documents with metadata_extraction_error set to None. Such documents are now processed normally.
Retrievers now handle top_k consistently. A negative top_k passed at runtime to InMemoryBM25Retriever, InMemoryEmbeddingRetriever or MultiRetriever (top_k and top_k_per_retriever) now raises a ValueError. Previously it was applied as a negative slice, silently dropping the last documents. A runtime top_k=0 returns no documents.
MultiRetriever now also validates top_k and top_k_per_retriever at initialization, raising a ValueError when they are set and not greater than 0, matching the in-memory retrievers. None still means no limit.
Fixed Pipeline.run(), Pipeline.run_async(), and Pipeline.stream() treating a flat input with a dictionary value as a component name. For example, {"payload": {"x": 1}} now reaches components with a payload input without requiring the component name in the input data.
Pipeline.run_async now lets a BreakpointException or PipelineRuntimeError raised by a nested component propagate unchanged, instead of wrapping it in another PipelineRuntimeError. This matches the synchronous Pipeline.run and preserves the original error context (for example an agent snapshot) when a component internally runs a pipeline.
XLSXToDocument now escapes pipe characters and replaces in-cell line breaks with spaces in the default Markdown pipe table output. This keeps cell content from being interpreted as additional table columns or rows.
@alanhuangyoo, @anakin87, @bilgeyucel, @carey-bk, @Cha-Imaa, @chrikrah, @dakjdakd, @gauravch-code, @Harsh23Kashyap, @Jayanth-reflex, @JbravoI, @julian-risch, @L4XB, @Lesereingrape, @lets-order-some-fries, @mnm-matin, @MohammadHijjawi97, @nanhanq1, @pcbeingused333, @PeterSmith0127-lcm, @Rainmemery, @sclfcz, @serhiizghama, @shivsin25, @ShousenZHANG, @simpleqt, @winklemad
anyio is installed transitively through httpx and openai ; versions before 4.14.2 are affected by CVE-2026-63374 ( GHSA-82r6-8w77-94w6 ).
SentenceWindowRetriever now queries the Document Store once per run or run_async call instead of once per retrieved document. Fewer calls mean lower latency and less load on the Document Store and outputs have not changed, so existing pipelines get faster without code changes.
InMemoryDocumentStore.bm25_retrieval and InMemoryBM25Retriever with BM25L (the default) or BM25Plus now return only documents that contain at least one query term, with or without scale_score. Previously, documents without any query term were returned with a positive score and filled up top_k. As a result, retrieval can now return fewer than top_k documents, or none at all. The delta lower bound now applies only to query terms that occur in the document, as in the original BM25L/BM25+ definitions. So scores of matching documents are lower than before whenever the query has terms that a document does not contain. If you filter results with a fixed score threshold, re-check the threshold. BM25Okapi is not affected.TextCleaner.run now raises a clear TypeError when texts is not a list or any element is not a str, instead of failing later or producing unexpected results.anyio>=4.14.2. anyio is installed transitively through httpx and openai; versions before 4.14.2 are affected by CVE-2026-63374 (GHSA-82r6-8w77-94w6).Fixed BM25 tokenization in InMemoryDocumentStore for Chinese, Japanese and Korean text. The default bm25_tokenization_regex now splits CJK characters into one token each, so bare-term queries can match words inside longer unspaced runs. Text is also NFC-normalized before tokenization, so composed and decomposed spellings produce the same tokens.
Fix Pipeline.connect() raising TypeError: object of type 'ellipsis' has no len() when one of the sockets is annotated with Callable[..., T]. A Callable[..., T] now matches callables with any parameters, in both directions, as long as the return types are compatible.
Fixed ChatPromptBuilder dropping content parts when the template is a list of ChatMessage objects. Only the first TextContent part was rendered, and every other part - additional texts, images, files -was removed from the rendered prompt without a warning. Template variables used in those dropped parts were not detected either, so they were not exposed as inputs. Now all TextContent parts are rendered and the remaining content parts are passed through unchanged.
Fixed CompactionHook.close() and close_async() to release the token counter's resources as well as the compactor's. Previously, resources such as OpenAITokenCounter's HTTP client remained open after the hook or Agent was closed.
ConfirmationHook no longer drops the messages that come after the last user or tool message when it rewrites the conversation, such as a system message added by another hook or an assistant answer followed by an on_exit reminder. Before, those messages were removed from the history the Agent sends to the model.
DocumentToImageContent no longer raises a KeyError for the whole batch when a document points to a PDF page that cannot be converted, because the page is out of range or the PDF cannot be read. That document now gets None in image_contents, like other invalid documents, and the rest of the batch is still converted. The logged warnings now include the path of the PDF file.
Fixed DOCXToDocument breaking a Markdown table when a cell contains a pipe or spans several paragraphs. A pipe was emitted as a column separator, and a cell's line break ended the row in the middle of it. Pipes are now escaped and line breaks inside a cell are collapsed to a space. The csv table format is unchanged: it already kept such cells intact by quoting them.
Fixed DOCXToDocument dropping hyperlink addresses inside tables. With link_format set to markdown or plain, links in table cells were written as their display text only, while links in body paragraphs kept their address. Links in table cells are now formatted the same way, in both the markdown and csv table formats. The default link_format="none" output is unchanged.
Fixed ConditionalRouter.from_dict mutating the caller's routes data in place: serialized output_type strings were deserialized directly inside the caller's dictionaries, so reusing the same serialized pipeline dict afterwards yielded already-deserialized type objects. from_dict now works on a copy and the caller's data is left untouched. BranchJoiner.from_dict received the same treatment.
LinkContentFetcher no longer adds its timeout and follow_redirects defaults to the client_kwargs dictionary passed by the caller. The dictionary is now copied before the defaults are applied, so reusing one HTTP client configuration across components no longer leaks these defaults.
Fix MetaFieldRanker to return no documents when missing_meta="drop" and all documents lack the ranking field or have a None value.
MetaFieldRanker now treats a Document whose meta_field value is None the same as a Document that is missing the field, applying the missing_meta setting to it. Previously a single None value made sorting fail, so the ranker logged a warning and returned all Documents in their original order, ignoring missing_meta="drop" as well.
MockTextEmbedder and MockDocumentEmbedder now accept non-positive dimension values when embedding or embedding_fn is provided, matching the documented behavior. The default deterministic embedding mode still requires a positive dimension.
ChatMessage.from_openai_dict_format now accepts tool-call arguments that are already a dictionary instead of raising a TypeError. Some OpenAI-compatible servers send a parsed object rather than a JSON string.
MetaFieldGroupingRanker now treats a group_by or subgroup_by value of None as missing, the same way it already treats sort_docs_by. Documents with None go to the end with the other ungrouped documents, instead of forming a group named "None" that also absorbed documents whose value was the string "None".
Fixed JSONConverter failing with ValueError when content_key contains numeric or boolean scalar values. These values are now converted to strings before creating the Document, while null values remain unchanged.
Fixed SentenceSplitter losing the whitespace between two sentences when the first one ends with a closing quote, as in He said "Hi." Bye.. Those characters were missing from the chunk text and shifted the split_idx_start offset of every following chunk, so chunks could no longer be mapped back onto the original text. This affects all components that split on sentences, such as DocumentSplitter, RecursiveDocumentSplitter, MarkdownHeaderSplitter and EmbeddingBasedDocumentSplitter. Chunk boundaries change for text that contains quoted sentences, so re-indexing an existing corpus produces different chunks than before.
Fixed LLMMetadataExtractor incorrectly treating a chat generator output that carries its own error field as a failed LLM call, which sent every document to failed_documents with metadata_extraction_error set to None. Such documents are now processed normally.
Retrievers now handle top_k consistently. A negative top_k passed at runtime to InMemoryBM25Retriever, InMemoryEmbeddingRetriever or MultiRetriever (top_k and top_k_per_retriever) now raises a ValueError. Previously it was applied as a negative slice, silently dropping the last documents. A runtime top_k=0 returns no documents.
MultiRetriever now also validates top_k and top_k_per_retriever at initialization, raising a ValueError when they are set and not greater than 0, matching the in-memory retrievers. None still means no limit.
Fixed Pipeline.run(), Pipeline.run_async(), and Pipeline.stream() treating a flat input with a dictionary value as a component name. For example, {"payload": {"x": 1}} now reaches components with a payload input without requiring the component name in the input data.
Pipeline.run_async now lets a BreakpointException or PipelineRuntimeError raised by a nested component propagate unchanged, instead of wrapping it in another PipelineRuntimeError. This matches the synchronous Pipeline.run and preserves the original error context (for example an agent snapshot) when a component internally runs a pipeline.
XLSXToDocument now escapes pipe characters and replaces in-cell line breaks with spaces in the default Markdown pipe table output. This keeps cell content from being interpreted as additional table columns or rows.
@alanhuangyoo, @anakin87, @bilgeyucel, @carey-bk, @Cha-Imaa, @chrikrah, @dakjdakd, @gauravch-code, @Harsh23Kashyap, @Jayanth-reflex, @JbravoI, @julian-risch, @L4XB, @Lesereingrape, @lets-order-some-fries, @mnm-matin, @MohammadHijjawi97, @nanhanq1, @pcbeingused333, @PeterSmith0127-lcm, @Rainmemery, @sclfcz, @serhiizghama, @shivsin25, @ShousenZHANG, @simpleqt, @winklemad
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
🧠 SummarizationCompactor: Compact Agent Context Without Losing It
The experimental SummarizationCompactor is a new compaction strategy for CompactionHook. It replaces older turns and Agent steps with LLM-generated summaries. Long-running Agents keep a condensed record of earlier goals, decisions, and open work. It summarizes as little as needed to reach the target size, starting with the oldest historical turns and only then moving to the current task's older steps. The min_keep_steps newest steps always stay as they are. You can use a small, cheap model to write the summaries.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.hooks.compaction import CompactionHook, SummarizationCompactor
compaction_hook = CompactionHook(
compactor=SummarizationCompactor(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-nano"),
min_keep_steps=2,
),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4"),
tools=[web_search],
hooks={"before_llm": [compaction_hook]},
)The experimental TokenBudgetHook caps how many tokens an Agent run can spend. Once the budget is used up, the Agent stops before its next LLM call, sets exit_reason to "token_budget_exceeded", and returns the messages collected so far. Set add_final_message=True to add an assistant message explaining why the run ended.
The hook is built on the new stop_run state key, which any hook can set to end a run cleanly with a custom reason instead of raising an error. For example, you can stop after too many calls to one tool:
from haystack.hooks import hook
from haystack.hooks.budget import TokenBudgetHook
@hook
def limit_searches(state):
if state.get("tool_call_counts", {}).get("web_search", 0) >= 5:
state.set("stop_run", "search_limit_reached")
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-5.4-nano"),
tools=[web_search],
hooks={"before_llm": [TokenBudgetHook(max_total_tokens=100_000), limit_searches]},
)
result = agent.run(messages)
print(result["exit_reason"]) # "text", "token_budget_exceeded", or "search_limit_reached"Building a pipeline takes less code with Haystack version 3.2: Pipeline.add_components() adds several components in one call and Pipeline.connect_many() creates several connections in one call. Both add_component() and add_components() now return the pipeline, so you can chain the calls.
from haystack import Pipeline
pipeline = (
Pipeline()
.add_components({"retriever": retriever, "prompt_builder": prompt_builder, "llm": llm})
.connect_many([("retriever", "prompt_builder.documents"), ("prompt_builder", "llm")])
)Serialized OutputAdapter and ConditionalRouter components containing Jinja custom_filters must now be loaded with Pipeline.load(..., unsafe=True) (or the equivalent Pipeline.loads / Pipeline.from_dict option).
The tool_result_offloaded meta key that the hook sets on an offloaded message now always holds a list of store references rather than a single reference string, since a result can span several entries. Code reading that key (for example to re-read an offloaded result) should index into the list.
ToolResultStore.write and read are typed str | bytes. Nothing changes at runtime for a text-only store, but one that subclasses the protocol and annotates content: str may need a type-checker fix.
The + operator for combining Toolsets has been removed. Pass Toolsets as a list wherever tools are accepted instead: Agent(..., tools=[toolset_a, toolset_b]). Pipelines serialized with a +-combined Toolset cannot be loaded anymore (their YAML references the removed internal _ToolsetWrapper class): recreate them with the list form and serialize them again.
Toolset.add() now only accepts a single Tool, not another Toolset. To combine Toolsets, pass them as a list: Agent(..., tools=[toolset_a, toolset_b]).
Add Pipeline.add_components() to add a mapping of component names to component instances in one call. Re-adding the same component instance under the same name with Pipeline.add_component() or Pipeline.add_components() is now a no-op.
pipeline.add_components(
{
"retriever": retriever,
"prompt_builder": prompt_builder,
"llm": llm,
}
)Added the experimental SummarizationCompactor (use with CompactionHook), which progressively summarizes a conversation until it fits a target token budget, preserving useful context from long-running Agents instead of dropping older messages.
It uses four summarization tiers in order until the target token budget is reached:
historical_turns: Starting with the oldest, summarize as few complete historical turns as needed.historical_summaries: When no complete historical turns remain, combine as few of the oldest historical summaries as needed.current_task_steps: Summarize the fewest oldest steps needed to reach the target while preserving the min_keep_steps newest steps.current_task_summaries: When no more steps can be summarized, combine as few of the current task's oldest summaries as needed to reach the target.from typing import Annotated
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.hooks.compaction import CompactionHook, SummarizationCompactor
from haystack.tools import tool
@tool
def web_search(query: Annotated[str, "The search query"]) -> str:
"""Search the web for current information."""
return f"Search results for: {query}"
agent_generator = OpenAIResponsesChatGenerator(model="gpt-5.4")
summary_generator = OpenAIResponsesChatGenerator(model="gpt-5.4-nano")
compaction_hook = CompactionHook(
compactor=SummarizationCompactor(
chat_generator=summary_generator,
min_keep_steps=2,
approximate_summary_tokens=1_024,
),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=agent_generator,
tools=[web_search],
hooks={"before_llm": [compaction_hook]},
)Added stop_run, a control flag that lets an Agent hook end a run cleanly instead of raising an error. The hook sets it to a reason of your choice (state.set("stop_run", "my_reason")), and the Agent stops without another model call and reports that reason in exit_reason.
Added TokenBudgetHook (experimental), a ready-made hook that caps how many tokens an Agent run may spend. The run ends with exit_reason set to "token_budget_exceeded", keeping the messages collected so far.
from haystack.hooks.budget import TokenBudgetHook
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-5.4-nano"),
tools=[search],
hooks={"before_llm": [TokenBudgetHook(max_total_tokens=100_000)]},
)Add Pipeline.connect_many() to connect multiple (sender, receiver) pairs in one call. Pipeline.add_component() and Pipeline.add_components() now return the pipeline, allowing pipeline-building methods to be chained.
pipeline = (
Pipeline()
.add_components(
{
"retriever": retriever,
"prompt_builder": prompt_builder,
"llm": llm,
}
)
.connect_many(
[
("retriever", "prompt_builder.documents"),
("prompt_builder", "llm"),
]
)
)Added the HAYSTACK_UNSAFE_DESERIALIZATION environment variable as a process-wide equivalent of loading with unsafe=True. When set to a truthy value (1 or true), every Pipeline.load / Pipeline.loads / Pipeline.from_dict call skips all deserialization safety checks — the module allowlist, the builtin/import-primitive and control-plane denylists, the object-internals traversal guard, and the refusal to honor a component's own unsafe: true flag. This is intended for deployments that only ever load fully trusted pipelines and cannot pass unsafe=True at every call site. A warning is logged the first time it takes effect. Only enable it when every pipeline the process loads is trusted: a single untrusted pipeline then leads to arbitrary code execution.
The switch is not limited to pipeline loading: it disables the same checks for every deserialization path in the process, including ones that take no unsafe argument of their own — Tool.from_dict, State.from_dict (agent snapshot resume), and the ConditionalRouter and OutputAdapter Jinja sandbox flags. A deployment that loads trusted pipelines at startup but accepts serialized tools or agent state at request time is therefore exposed on the request path as well, not only at load time.
The variable is read once, on the first deserialization in the process, and the result is then frozen for the process lifetime: writes to it afterwards are ignored, in either direction. Set it before the first pipeline is loaded. Freezing keeps the safety mode of a process from changing under a caller's feet and, because the first read happens before any deserialized data can run, stops a hostile pipeline from switching the checks off while it is being loaded.
FileSystemSkillStore's skills_dir now also accepts a Secret, so the skills directory path can be sourced from an environment variable (e.g. Secret.from_env_var("SKILLS_DIR")) instead of being hard-coded in pipeline configuration.
AnswerJoiner now logs an informational message when sort_by_score is enabled and some answers have no score, matching the existing behaviour of DocumentJoiner. Such answers are still sorted as if their score were -infinity, but the demotion is no longer silent. This is common when joining GeneratedAnswer objects, which carry no score at all.
Add the min_content_length parameter to DocumentCleaner. After the configured cleaning steps run, text documents shorter than the threshold (ignoring leading and trailing whitespace) are omitted from the output.
Add split_by="token" to DocumentSplitter and DocumentPreprocessor. Splits text by LLM token count using tiktoken. The encoding defaults to "o200k_base" (current OpenAI models) and can be changed via the new tokenizer_encoding parameter. split_length, split_overlap, and split_threshold all work as usual. Requires pip install tiktoken.
ToolResultOffloadHook now offloads tool results carrying ImageContent or FileContent, not just text.
Each part of a result gets its own store entry, with images and files decoded from base64 to raw bytes, under the extension of its filename when it has one and of its mime_type otherwise (falling back to .bin). For example:
Tool result offloaded to 3 files:
1. text (412 characters) at '/abs/path/tool_results/2_fetch_call-123_0.txt'. Preview: Quarterly report...
2. image/png (48210 bytes) at '/abs/path/tool_results/2_fetch_call-123_1.png'
3. application/pdf named 'q3.pdf' (1048576 bytes) at '/abs/path/tool_results/2_fetch_call-123_2.pdf'
Binary content is opt-in, via a new supports_binary_content attribute on the ToolResultStore protocol, which FileSystemToolResultStore sets to True. A store leaving it at its False default never receives bytes: images and files stay in context with a warning, as before, so existing custom stores keep working unchanged.
An offload policy receives the result as a single string: the text and base64 payloads of all its parts concatenated.
FallbackChatGenerator now wraps every internal ChatGenerator attempt in a haystack.chat_generator.run tracing span for both synchronous and asynchronous execution. With content tracing enabled, each internal span records the forwarded inputs. The successful output remains on the enclosing FallbackChatGenerator component span so tracing backends do not count its token usage twice.
TopPSampler now rejects negative, boolean, and non-integer min_top_k values during initialization with a clear ValueError.
Pipeline component outputs typed as list[T] to connect to inputs typed as Iterable[T]. Lists already satisfy the iterable input contract and are passed through without conversion.AnswerJoiner.run docstring: it did not mention that the sort_by_score init parameter affects the order of the returned answers. The docstring now documents that when sort_by_score=True the merged answers are sorted by score in descending order (answers without a score are handled as if their score were -infinity) before top_k is applied, and that the input order is preserved otherwise.Agent.clone() and Agent.to_dict()/from_dict() raising TypeError for an Agent built on a chat generator that does not accept a tools parameter. tools is normalized to an empty list at init, and both round-trip paths passed that empty list back to the constructor, where it was treated as "tools were provided". An empty list now carries no tools, matching the equivalent check already applied in Agent.run(). Passing tools=[] explicitly to such an Agent is accepted for the same reason. A non-empty tools value still raises, and a Toolset is still never tested for truthiness at init.Agent from retrying an empty response when a chat generator reports finish_reason="length". This can happen when reasoning tokens exhaust the maximum output token budget. OpenAIResponsesChatGenerator now maps terminal Responses API events to Haystack FinishReason values and attaches finish_reason to the returned ChatMessage.Agent now reports exit_reason="length" or "content_filter" when a tool-call-free model reply ends for either reason, including replies containing partial text. These incomplete generations stop by default but can be recovered with an on_exit hook. AgentTool also warns the calling model that such results may be incomplete.JsonSchemaValidator now accepts an empty JSON Schema ({}) and honors it when passed to run(). Empty schemas are valid JSON Schema and match every JSON value.AutoMergingRetriever rejecting documents whose hierarchy metadata uses __level=0 or __block_size=0. HierarchicalDocumentSplitter stores those values on the root node, so validation now checks for key presence instead of truthiness.str instead of repr, because some carry their whole input in their repr - UnicodeDecodeError keeps the entire buffer it failed to decode, so logging one raised while decoding a large file would emit that whole file as a single log line. Any remaining value longer than 4096 characters is truncated.LLMMetadataExtractor now removes the metadata_extraction_error and metadata_extraction_response keys from a document's metadata on every successful extraction, in both run and run_async. Previously, when a document from failed_documents was re-run and the LLM returned an empty JSON object {}, the document ended up in documents still carrying both keys from the earlier failed attempt.TextFileToDocument, CSVToDocument, MarkdownToDocument and MultiFileConverter now default to utf-8-sig instead of utf-8, so a UTF-8 byte order mark (BOM) is stripped rather than leaking into Document.content as a zero-width character. utf-8-sig decodes plain UTF-8 identically, so files without a BOM are unaffected.JSONConverter no longer silently skips UTF-8 files that start with a byte order mark. The file content is now decoded with utf-8-sig. Previously the UnicodeError was caught and logged as a warning, so a BOM file was dropped from the pipeline without raising - three input files could produce two documents with no error.DocumentJoiner now raises a ValueError when any value in weights is negative. Previously negative weights were normalized by their -- possibly negative -- sum, which flipped their sign and could produce negative document scores in merge mode. For example weights=[1, -2] was normalized to [-1.0, 2.0]. The existing error for weights that sum to zero is unchanged.DocumentJoiner distribution-based rank fusion when an input list contains a single document or has zero score variance.DocumentSplitter now adds split_id, split_idx_start and page_number to the metadata of the chunks it creates when split_by="function", like it already did for all other split modes. Without this metadata, components that rely on it, such as SentenceWindowRetriever, could not be used with chunks produced by a custom splitting function. Empty chunks returned by the splitting function are now also skipped, unless skip_empty_documents=False.DocumentSplitter with split_by="token" generating redundant trailing chunks containing only overlap. The splitter now stops creating chunks after reaching the end of the document.DocumentToImageContent no longer raises a ValueError for the whole batch when one document is missing the file path or page_number metadata, has an invalid file path, or has an unsupported MIME type. That document now gets None in image_contents and a warning with the reason is logged, while the other documents are still converted. This lets LLMDocumentContentExtractor return such documents in failed_documents instead of failing the entire run.ToolResultOffloadHook offloading empty tool results. An empty string, or a result made up of nothing but empty text blocks, was written to the ToolResultStore as a zero-byte entry and replaced in the conversation by a pointer ending in a dangling Preview: - growing the context the hook exists to shrink, and inviting the model to read back an empty file. Such results are now left in context, as empty-list results already were.FilterRetriever so that an empty runtime filters dictionary clears filters provided at initialization during synchronous and asynchronous execution.document_matches_filter raising AttributeError for dotted filter fields whose root is not a Document attribute (for example {"field": "typo.x", ...}). Such fields are now treated as missing, consistent with non-dotted unknown fields and the documented behavior. This affected InMemoryDocumentStore.filter_documents and MetadataRouter.CompactionHook mis-estimating the context size when compaction leaves a conversation with no assistant message, for example when a compactor summarizes every Agent step away. The hook now records the whole compacted conversation as accounted for, and the estimate reads that count back unchanged, so a second compaction hook running right after sees the true size instead of counting the conversation twice.AzureOpenAIChatGenerator.to_dict() when response_format is passed as a dictionary.ComponentDevice.first_device raising a ValueError when the first entry of a device map is a disk device. Disk devices are only valid as part of a device map (for example when HuggingFace offloads weights with device_map='auto'), so first_device now skips disk entries and returns the first usable device instead of crashing. If the device map is empty or contains only disk devices, a ValueError is raised.deserialize_secrets_inplace with recursive=True: serialized secrets were left as plain dicts instead of being converted back to Secret, and nested dictionaries deeper than one level were never visited.FileTypeRouter.run() writing the meta passed to it into the meta dict of the input ByteStream objects. The metadata is now added to a copy, so ByteStream sources are left untouched, matching the behaviour already in place for file path sources.FilterPolicy.MERGE from mutating initialization or runtime filters when combining logical filters. Reusing a retriever for multiple runs now applies each runtime filter independently.InMemoryDocumentStore.count_unique_metadata_by_filter and count_unique_metadata_by_filter_async raising TypeError when metadata values are JSON-serializable lists or dictionaries, and make composite metadata deduplication consistent with get_metadata_field_unique_values.JsonSchemaValidator no longer raises ValueError when the message content is a top-level JSON scalar like "hello", 42, true or null. The same crash happened for JSON arrays of scalars. Those values now reach the schema validator, which either accepts them or returns the usual validation_error output.LinkContentFetcher.run_async() omitting default headers configured through client_kwargs. Asynchronous and synchronous fetches now apply the same client-default headers.DocumentSplitter and RecursiveDocumentSplitter failing on documents containing strings such as <|endoftext|>. These strings are now preserved and counted as ordinary text, including when splitting with overlap.LLMDocumentContentExtractor.run_async converting documents to images on the event loop. DocumentToImageContent reads every file and renders the requested PDF pages, and it has no run_async, so calling it directly blocked the loop for the whole batch before any LLM call started. It now runs through _execute_component_async, which hands the synchronous work to a thread.parent_headers metadata in MarkdownHeaderSplitter when keep_headers=False and secondary splitting is enabled. Editing one chunk's parent headers no longer changes its siblings' metadata.MetaFieldGroupingRanker no longer partially reorders a group before falling back on a TypeError. When sort_docs_by values are mutually non-comparable (e.g. int and str), the group's original insertion order is now guaranteed, as the sorting happens on a copy instead of in place.AssertionError crash when comparing a Pipeline to any non-Pipeline object (such as pipeline == object() or pipeline in [object(), ...]) caused by an inverted isinstance check in PipelineBase.__eq__.QueryExpander now returns an empty query list when query is None or not a string, instead of raising AttributeError on str.strip.RecursiveDocumentSplitter applied the overlap at every recursion level when chunking text with multiple separators, producing chunks that were not substrings of the source text. The overlap is now applied exactly once, on the final chunk list.Tool subclasses such as ComponentTool in OpenAIResponsesChatGenerator.from_dict() and AzureOpenAIResponsesChatGenerator.from_dict().FileSystemSkillStore.load_skill including symlinks that resolve outside the skill directory in its bundled-file manifest. Escaping symlinks are now excluded, matching read_skill_file traversal protection.split_overlap validation message in RecursiveDocumentSplitter (0 is the default and only negative values are rejected) and clarify that overlap is measured in split_units. Also fix stale docstrings in AnswerJoiner (documented parameters the methods do not take, and claimed sorting that only happens for answers) and the meta_field ranker (referenced a nonexistent score mode).create_tool_from_function, the @tool decorator, and ComponentTool stripping title keys inside OpenAPI 3.0 singular example values when building a tool's JSON schema. example is now treated as instance data alongside default, const, enum and examples, so a title key carried via Pydantic json_schema_extra stays part of the tool contract.MetaFieldRanker raising a TypeError when meta_value_type is set and a document contains an unhashable metadata value, such as a list or dictionary. These values now follow the existing warning and original-document fallback behavior.XLSXToDocument skipping workbooks with hyperlinks in numeric or boolean columns when using pandas 3. Hyperlink text is now inserted without a dtype error, while columns without hyperlinks retain their original types and formatting.LinkContentFetcher now performs the configured number of retries in synchronous runs. Previously, retry_attempts counted the initial request as an attempt, so synchronous runs performed one fewer retry than configured and behaved differently from asynchronous runs.LLMDocumentContentExtractor incorrectly treating a chat generator output that carries its own error field as a failed LLM call. Such documents are now processed normally.LLMMessagesRouter now raises a ValueError when output_names contains "chat_generator_text" or "unmatched". Those two outputs are always created by the router, so reusing either name used to be accepted and then silently broke the component: "chat_generator_text" had its socket retyped to list[ChatMessage] and lost the LLM reply at run time, while "unmatched" made a matched decision indistinguishable from an unmatched one.LLMMetadataExtractor incorrectly treating an extracted error metadata field as an extraction failure. error can now be used in expected_keys like any other metadata key.LLMRanker no longer crashes with AttributeError when query is None (or another non-string) coming from an upstream pipeline component. Both run and run_async now return the documents unranked, matching the existing empty-query fallback.MarkdownHeaderSplitter silently dropping a document's leading header line when secondary_split is used with keep_headers=False. This happened for any chunk that wasn't the result of an actual header split, for example when a document has no headers, only headers without body content, a header at a level excluded from header_split_levels, or input metadata that already contained a header key. The header line is now only stripped from chunks that truly came from a header split.MarkdownHeaderSplitter no longer drops the text that precedes the first header it splits on. A document's opening paragraph, title block, or front matter was silently lost, and so was any content above the first matching header when header_split_levels excluded the levels above it. A document whose headers all had empty bodies lost that text too, because it was treated as header-only. That text is now emitted as a leading chunk with empty header and parent_headers metadata, so every chunk still carries the documented metadata fields. Leading text that is only whitespace joins the first chunk instead of becoming a chunk of its own, which keeps the chunks a byte-exact partition of the input without adding an empty chunk.page_number metadata from MarkdownHeaderSplitter when secondary splitting uses overlap. Page breaks in overlapping content are no longer counted multiple times, and custom page_break_character values are respected.EmbeddingBasedDocumentSplitter emitting a final split shorter than min_length. Small splits were only merged forward, so the last one had nothing left to absorb and was returned as its own document. It is now merged into the preceding split, unless doing so would reach max_length, the same limit that already governs forward merges.MultiQueryTextRetriever, MultiQueryEmbeddingRetriever and MultiRetriever now honour max_workers in their run_async methods. Previously only the synchronous run bounded fan-out (via a ThreadPoolExecutor); run_async launched every sub-query / sub-retriever call at once, ignoring max_workers and risking rate limits or connection exhaustion against the underlying retriever or embedder. run_async now bounds concurrency with an asyncio.Semaphore, matching the behaviour of the LLM extractors.ChatMessage.from_openai_dict_format accepts the same empty content, so such a message round-trips.OpenAIDocumentEmbedder.run_async (and AzureOpenAIDocumentEmbedder) now request encoding_format="float" from the embeddings endpoint, matching the synchronous run. The async batch path was missed when this was added in #9655, so run_async let the OpenAI SDK negotiate the base64 wire format, which some OpenAI-compatible endpoints do not support.OpenAIChatGenerator and OpenAIResponsesChatGenerator so that an empty runtime tools list overrides tools configured at initialization.timeout and max_retries to the to_dict method of OpenAIImageGenerator. They were dropped on serialization, so a pipeline saved and reloaded silently fell back to the OPENAI_TIMEOUT/OPENAI_MAX_RETRIES defaults instead of the configured values. This matches OpenAIChatGenerator, OpenAITextEmbedder and OpenAIDocumentEmbedder, which already serialize both.OutputAdapter and ConditionalRouter no longer coerce string outputs to other Python types when output_type=str. Previously the rendered template was always passed through ast.literal_eval in safe mode, so a string that happened to be a valid Python literal was silently converted (for example "1,000" became the tuple (1, 0) and "42" became the integer 42), producing a value that did not match the declared str output type. Literal evaluation is now skipped when output_type is str, so string outputs are returned unchanged. Other output types still reconstruct structured literals as before.PythonCodeSplitter now deep-copies the metadata of the document it splits, matching the other splitters. Previously it copied it shallowly, so nested values such as a list under meta["tags"] were shared between every chunk and with the input document, and editing one chunk's metadata changed all the others. The secondary line-based split of oversized units had the same problem.RecursiveDocumentSplitter and TextCleaner now keep their init parameters when serialized, for example in a pipeline saved with Pipeline.dumps(). Before, RecursiveDocumentSplitter wrote split_unit: word whatever unit it was created with, so a reloaded splitter configured for characters or tokens counted words instead. A reloaded TextCleaner lost all of its options and returned the texts unchanged.DocumentSplitter with split_by="word", "period", "page", "passage", "line" or "sentence" generating redundant trailing chunks containing only overlap. Segments that add no new text beyond the already covered units are now skipped, mirroring the split_by="token" fix.IndexError: list index out of range when a chat generator's stream completes without emitting any chunks. _convert_streaming_chunks_to_chat_message now returns an assistant message with empty content and None metadata instead.SuperComponent.run_async now returns outputs mapped from components that are not leaves of the wrapped pipeline, like run does. Before, an output that was also consumed inside the pipeline, such as retriever.documents feeding a prompt builder, was silently missing from the async result. This also affected a PipelineTool run by Agent.run_async, whose outputs_to_state got nothing for such outputs.create_tool_from_function and the @tool decorator now work with functions defined in a module that uses from __future__ import annotations. Before, the parameter annotations were plain strings, and creating the tool raised a SchemaGenerationError for any parameter typed with something other than a builtin, such as Annotated, Literal, Optional, Document or State. The annotations are now resolved first, and Annotated descriptions are kept.MetadataRouter now raises a clear ValueError when a rule uses the reserved output name unmatched, instead of failing with an internal duplicate-keyword TypeError during output registration.FileTypeRouter and DocumentTypeRouter now raise a clear ValueError when a MIME type uses a reserved output name (unclassified, or failed on FileTypeRouter), instead of failing with an internal duplicate-keyword TypeError during output registration.XLSXToDocument writing the string nan into an empty cell when table_format="markdown". The same cell is written as an empty field by table_format="csv", so an empty cell now reads as empty in both formats. A different placeholder can be set with table_format_kwargs={"missingval": "N/A"}.@abhati27, @abo3losh1, @Aftabbs, @AjayShivran, @alanhuangyoo, @anakin87, @Anurag-M1, @Arman-Beykmohammadi, @ArzelaAscoIi, @atikulmunna, @Awshesh12, @bilgeyucel, @bogdankostic, @businessarshgoyal, @coder058, @CoralGarden52, @davidsbatista, @dfedoryshchev, @dingpuyu, @Diwak4r, @Dxfory, @eatLaoJun, @ege-arhan, @feiiiiii5, @fng713, @Goodnight77, @Gout999, @gyanu2507, @Harsh23Kashyap, @iridescentWen, @JbravoI, @JimmyWang0417, @jliounis, @julian-risch, @kacperlukawski, @Koushik890, @Kuang-xianxin, @L4XB, @Lesereingrape, @lets-order-some-fries, @LHMQ878, @linhongyu510, @mfurkanakinci, @mikemikimike, @MohammadHijjawi97, @mrchtr, @Nikhi00718, @otiscuilei, @PattonBrown, @pcbeingused333, @PeterSmith0127-lcm, @Rainmemery, @rautaditya2606, @Ricky-7-Yan, @sainikhiljuluri, @samrusani, @seanxuu, @ShousenZHANG, @simpleqt, @sjrl, @spacesheepinternet, @SyedShahmeerAli12, @teachershuang, @tstadel, @uczltw6, @Vedant-Agarwal, @vercel[bot], @winter-street, @xblwh, @yavuz-yilmaz
Serialized OutputAdapter and ConditionalRouter components containing Jinja custom_filters must now be loaded with Pipeline.load(..., unsafe=True) (or
Serialized OutputAdapter and ConditionalRouter components containing Jinja custom_filters must now be loaded with Pipeline.load(..., unsafe=True) (or the equivalent Pipeline.loads / Pipeline.from_dict option).
The tool_result_offloaded meta key that the hook sets on an offloaded message now always holds a list of store references rather than a single reference string, since a result can span several entries. Code reading that key (for example to re-read an offloaded result) should index into the list.
ToolResultStore.write and read are typed str | bytes. Nothing changes at runtime for a text-only store, but one that subclasses the protocol and annotates content: str may need a type-checker fix.
The + operator for combining Toolsets has been removed. Pass Toolsets as a list wherever tools are accepted instead: Agent(..., tools=[toolset_a, toolset_b]). Pipelines serialized with a +-combined Toolset cannot be loaded anymore (their YAML references the removed internal _ToolsetWrapper class): recreate them with the list form and serialize them again.
Toolset.add() now only accepts a single Tool, not another Toolset. To combine Toolsets, pass them as a list: Agent(..., tools=[toolset_a, toolset_b]).
Add Pipeline.add_components() to add a mapping of component names to component instances in one call. Re-adding the same component instance under the same name with Pipeline.add_component() or Pipeline.add_components() is now a no-op.
pipeline.add_components(
{
"retriever": retriever,
"prompt_builder": prompt_builder,
"llm": llm,
}
)Added the experimental SummarizationCompactor (use with CompactionHook), which progressively summarizes a conversation until it fits a target token budget, preserving useful context from long-running Agents instead of dropping older messages.
It uses four summarization tiers in order until the target token budget is reached:
historical_turns: Starting with the oldest, summarize as few complete historical turns as needed.historical_summaries: When no complete historical turns remain, combine as few of the oldest historical summaries as needed.current_task_steps: Summarize the fewest oldest steps needed to reach the target while preserving the min_keep_steps newest steps.current_task_summaries: When no more steps can be summarized, combine as few of the current task's oldest summaries as needed to reach the target.from typing import Annotated
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.hooks.compaction import CompactionHook, SummarizationCompactor
from haystack.tools import tool
@tool
def web_search(query: Annotated[str, "The search query"]) -> str:
"""Search the web for current information."""
return f"Search results for: {query}"
agent_generator = OpenAIResponsesChatGenerator(model="gpt-5.4")
summary_generator = OpenAIResponsesChatGenerator(model="gpt-5.4-nano")
compaction_hook = CompactionHook(
compactor=SummarizationCompactor(
chat_generator=summary_generator,
min_keep_steps=2,
approximate_summary_tokens=1_024,
),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=agent_generator,
tools=[web_search],
hooks={"before_llm": [compaction_hook]},
)Added stop_run, a control flag that lets an Agent hook end a run cleanly instead of raising an error. The hook sets it to a reason of your choice (state.set("stop_run", "my_reason")), and the Agent stops without another model call and reports that reason in exit_reason.
Added TokenBudgetHook (experimental), a ready-made hook that caps how many tokens an Agent run may spend. The run ends with exit_reason set to "token_budget_exceeded", keeping the messages collected so far.
from haystack.hooks.budget import TokenBudgetHook
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-5.4-nano"),
tools=[search],
hooks={"before_llm": [TokenBudgetHook(max_total_tokens=100_000)]},
)Add Pipeline.connect_many() to connect multiple (sender, receiver) pairs in one call. Pipeline.add_component() and Pipeline.add_components() now return the pipeline, allowing pipeline-building methods to be chained.
pipeline = (
Pipeline()
.add_components(
{
"retriever": retriever,
"prompt_builder": prompt_builder,
"llm": llm,
}
)
.connect_many(
[
("retriever", "prompt_builder.documents"),
("prompt_builder", "llm"),
]
)
)Added the HAYSTACK_UNSAFE_DESERIALIZATION environment variable as a process-wide equivalent of loading with unsafe=True. When set to a truthy value (1 or true), every Pipeline.load / Pipeline.loads / Pipeline.from_dict call skips all deserialization safety checks — the module allowlist, the builtin/import-primitive and control-plane denylists, the object-internals traversal guard, and the refusal to honor a component's own unsafe: true flag. This is intended for deployments that only ever load fully trusted pipelines and cannot pass unsafe=True at every call site. A warning is logged the first time it takes effect. Only enable it when every pipeline the process loads is trusted: a single untrusted pipeline then leads to arbitrary code execution.
The switch is not limited to pipeline loading: it disables the same checks for every deserialization path in the process, including ones that take no unsafe argument of their own — Tool.from_dict, State.from_dict (agent snapshot resume), and the ConditionalRouter and OutputAdapter Jinja sandbox flags. A deployment that loads trusted pipelines at startup but accepts serialized tools or agent state at request time is therefore exposed on the request path as well, not only at load time.
The variable is read once, on the first deserialization in the process, and the result is then frozen for the process lifetime: writes to it afterwards are ignored, in either direction. Set it before the first pipeline is loaded. Freezing keeps the safety mode of a process from changing under a caller's feet and, because the first read happens before any deserialized data can run, stops a hostile pipeline from switching the checks off while it is being loaded.
FileSystemSkillStore's skills_dir now also accepts a Secret, so the skills directory path can be sourced from an environment variable (e.g. Secret.from_env_var("SKILLS_DIR")) instead of being hard-coded in pipeline configuration.
AnswerJoiner now logs an informational message when sort_by_score is enabled and some answers have no score, matching the existing behaviour of DocumentJoiner. Such answers are still sorted as if their score were -infinity, but the demotion is no longer silent. This is common when joining GeneratedAnswer objects, which carry no score at all.
Add the min_content_length parameter to DocumentCleaner. After the configured cleaning steps run, text documents shorter than the threshold (ignoring leading and trailing whitespace) are omitted from the output.
Add split_by="token" to DocumentSplitter and DocumentPreprocessor. Splits text by LLM token count using tiktoken. The encoding defaults to "o200k_base" (current OpenAI models) and can be changed via the new tokenizer_encoding parameter. split_length, split_overlap, and split_threshold all work as usual. Requires pip install tiktoken.
ToolResultOffloadHook now offloads tool results carrying ImageContent or FileContent, not just text.
Each part of a result gets its own store entry, with images and files decoded from base64 to raw bytes, under the extension of its filename when it has one and of its mime_type otherwise (falling back to .bin). For example:
Tool result offloaded to 3 files:
1. text (412 characters) at '/abs/path/tool_results/2_fetch_call-123_0.txt'. Preview: Quarterly report...
2. image/png (48210 bytes) at '/abs/path/tool_results/2_fetch_call-123_1.png'
3. application/pdf named 'q3.pdf' (1048576 bytes) at '/abs/path/tool_results/2_fetch_call-123_2.pdf'
Binary content is opt-in, via a new supports_binary_content attribute on the ToolResultStore protocol, which FileSystemToolResultStore sets to True. A store leaving it at its False default never receives bytes: images and files stay in context with a warning, as before, so existing custom stores keep working unchanged.
An offload policy receives the result as a single string: the text and base64 payloads of all its parts concatenated.
FallbackChatGenerator now wraps every internal ChatGenerator attempt in a haystack.chat_generator.run tracing span for both synchronous and asynchronous execution. With content tracing enabled, each internal span records the forwarded inputs. The successful output remains on the enclosing FallbackChatGenerator component span so tracing backends do not count its token usage twice.
TopPSampler now rejects negative, boolean, and non-integer min_top_k values during initialization with a clear ValueError.
Pipeline component outputs typed as list[T] to connect to inputs typed as Iterable[T]. Lists already satisfy the iterable input contract and are passed through without conversion.AnswerJoiner.run docstring: it did not mention that the sort_by_score init parameter affects the order of the returned answers. The docstring now documents that when sort_by_score=True the merged answers are sorted by score in descending order (answers without a score are handled as if their score were -infinity) before top_k is applied, and that the input order is preserved otherwise.Agent.clone() and Agent.to_dict()/from_dict() raising TypeError for an Agent built on a chat generator that does not accept a tools parameter. tools is normalized to an empty list at init, and both round-trip paths passed that empty list back to the constructor, where it was treated as "tools were provided". An empty list now carries no tools, matching the equivalent check already applied in Agent.run(). Passing tools=[] explicitly to such an Agent is accepted for the same reason. A non-empty tools value still raises, and a Toolset is still never tested for truthiness at init.Agent from retrying an empty response when a chat generator reports finish_reason="length". This can happen when reasoning tokens exhaust the maximum output token budget. OpenAIResponsesChatGenerator now maps terminal Responses API events to Haystack FinishReason values and attaches finish_reason to the returned ChatMessage.Agent now reports exit_reason="length" or "content_filter" when a tool-call-free model reply ends for either reason, including replies containing partial text. These incomplete generations stop by default but can be recovered with an on_exit hook. AgentTool also warns the calling model that such results may be incomplete.JsonSchemaValidator now accepts an empty JSON Schema ({}) and honors it when passed to run(). Empty schemas are valid JSON Schema and match every JSON value.AutoMergingRetriever rejecting documents whose hierarchy metadata uses __level=0 or __block_size=0. HierarchicalDocumentSplitter stores those values on the root node, so validation now checks for key presence instead of truthiness.str instead of repr, because some carry their whole input in their repr - UnicodeDecodeError keeps the entire buffer it failed to decode, so logging one raised while decoding a large file would emit that whole file as a single log line. Any remaining value longer than 4096 characters is truncated.LLMMetadataExtractor now removes the metadata_extraction_error and metadata_extraction_response keys from a document's metadata on every successful extraction, in both run and run_async. Previously, when a document from failed_documents was re-run and the LLM returned an empty JSON object {}, the document ended up in documents still carrying both keys from the earlier failed attempt.TextFileToDocument, CSVToDocument, MarkdownToDocument and MultiFileConverter now default to utf-8-sig instead of utf-8, so a UTF-8 byte order mark (BOM) is stripped rather than leaking into Document.content as a zero-width character. utf-8-sig decodes plain UTF-8 identically, so files without a BOM are unaffected.JSONConverter no longer silently skips UTF-8 files that start with a byte order mark. The file content is now decoded with utf-8-sig. Previously the UnicodeError was caught and logged as a warning, so a BOM file was dropped from the pipeline without raising - three input files could produce two documents with no error.DocumentJoiner now raises a ValueError when any value in weights is negative. Previously negative weights were normalized by their -- possibly negative -- sum, which flipped their sign and could produce negative document scores in merge mode. For example weights=[1, -2] was normalized to [-1.0, 2.0]. The existing error for weights that sum to zero is unchanged.DocumentJoiner distribution-based rank fusion when an input list contains a single document or has zero score variance.DocumentSplitter now adds split_id, split_idx_start and page_number to the metadata of the chunks it creates when split_by="function", like it already did for all other split modes. Without this metadata, components that rely on it, such as SentenceWindowRetriever, could not be used with chunks produced by a custom splitting function. Empty chunks returned by the splitting function are now also skipped, unless skip_empty_documents=False.DocumentSplitter with split_by="token" generating redundant trailing chunks containing only overlap. The splitter now stops creating chunks after reaching the end of the document.DocumentToImageContent no longer raises a ValueError for the whole batch when one document is missing the file path or page_number metadata, has an invalid file path, or has an unsupported MIME type. That document now gets None in image_contents and a warning with the reason is logged, while the other documents are still converted. This lets LLMDocumentContentExtractor return such documents in failed_documents instead of failing the entire run.ToolResultOffloadHook offloading empty tool results. An empty string, or a result made up of nothing but empty text blocks, was written to the ToolResultStore as a zero-byte entry and replaced in the conversation by a pointer ending in a dangling Preview: - growing the context the hook exists to shrink, and inviting the model to read back an empty file. Such results are now left in context, as empty-list results already were.FilterRetriever so that an empty runtime filters dictionary clears filters provided at initialization during synchronous and asynchronous execution.document_matches_filter raising AttributeError for dotted filter fields whose root is not a Document attribute (for example {"field": "typo.x", ...}). Such fields are now treated as missing, consistent with non-dotted unknown fields and the documented behavior. This affected InMemoryDocumentStore.filter_documents and MetadataRouter.CompactionHook mis-estimating the context size when compaction leaves a conversation with no assistant message, for example when a compactor summarizes every Agent step away. The hook now records the whole compacted conversation as accounted for, and the estimate reads that count back unchanged, so a second compaction hook running right after sees the true size instead of counting the conversation twice.AzureOpenAIChatGenerator.to_dict() when response_format is passed as a dictionary.ComponentDevice.first_device raising a ValueError when the first entry of a device map is a disk device. Disk devices are only valid as part of a device map (for example when HuggingFace offloads weights with device_map='auto'), so first_device now skips disk entries and returns the first usable device instead of crashing. If the device map is empty or contains only disk devices, a ValueError is raised.deserialize_secrets_inplace with recursive=True: serialized secrets were left as plain dicts instead of being converted back to Secret, and nested dictionaries deeper than one level were never visited.FileTypeRouter.run() writing the meta passed to it into the meta dict of the input ByteStream objects. The metadata is now added to a copy, so ByteStream sources are left untouched, matching the behaviour already in place for file path sources.FilterPolicy.MERGE from mutating initialization or runtime filters when combining logical filters. Reusing a retriever for multiple runs now applies each runtime filter independently.InMemoryDocumentStore.count_unique_metadata_by_filter and count_unique_metadata_by_filter_async raising TypeError when metadata values are JSON-serializable lists or dictionaries, and make composite metadata deduplication consistent with get_metadata_field_unique_values.JsonSchemaValidator no longer raises ValueError when the message content is a top-level JSON scalar like "hello", 42, true or null. The same crash happened for JSON arrays of scalars. Those values now reach the schema validator, which either accepts them or returns the usual validation_error output.LinkContentFetcher.run_async() omitting default headers configured through client_kwargs. Asynchronous and synchronous fetches now apply the same client-default headers.DocumentSplitter and RecursiveDocumentSplitter failing on documents containing strings such as <|endoftext|>. These strings are now preserved and counted as ordinary text, including when splitting with overlap.LLMDocumentContentExtractor.run_async converting documents to images on the event loop. DocumentToImageContent reads every file and renders the requested PDF pages, and it has no run_async, so calling it directly blocked the loop for the whole batch before any LLM call started. It now runs through _execute_component_async, which hands the synchronous work to a thread.parent_headers metadata in MarkdownHeaderSplitter when keep_headers=False and secondary splitting is enabled. Editing one chunk's parent headers no longer changes its siblings' metadata.MetaFieldGroupingRanker no longer partially reorders a group before falling back on a TypeError. When sort_docs_by values are mutually non-comparable (e.g. int and str), the group's original insertion order is now guaranteed, as the sorting happens on a copy instead of in place.AssertionError crash when comparing a Pipeline to any non-Pipeline object (such as pipeline == object() or pipeline in [object(), ...]) caused by an inverted isinstance check in PipelineBase.__eq__.QueryExpander now returns an empty query list when query is None or not a string, instead of raising AttributeError on str.strip.RecursiveDocumentSplitter applied the overlap at every recursion level when chunking text with multiple separators, producing chunks that were not substrings of the source text. The overlap is now applied exactly once, on the final chunk list.Tool subclasses such as ComponentTool in OpenAIResponsesChatGenerator.from_dict() and AzureOpenAIResponsesChatGenerator.from_dict().FileSystemSkillStore.load_skill including symlinks that resolve outside the skill directory in its bundled-file manifest. Escaping symlinks are now excluded, matching read_skill_file traversal protection.split_overlap validation message in RecursiveDocumentSplitter (0 is the default and only negative values are rejected) and clarify that overlap is measured in split_units. Also fix stale docstrings in AnswerJoiner (documented parameters the methods do not take, and claimed sorting that only happens for answers) and the meta_field ranker (referenced a nonexistent score mode).create_tool_from_function, the @tool decorator, and ComponentTool stripping title keys inside OpenAPI 3.0 singular example values when building a tool's JSON schema. example is now treated as instance data alongside default, const, enum and examples, so a title key carried via Pydantic json_schema_extra stays part of the tool contract.MetaFieldRanker raising a TypeError when meta_value_type is set and a document contains an unhashable metadata value, such as a list or dictionary. These values now follow the existing warning and original-document fallback behavior.XLSXToDocument skipping workbooks with hyperlinks in numeric or boolean columns when using pandas 3. Hyperlink text is now inserted without a dtype error, while columns without hyperlinks retain their original types and formatting.LinkContentFetcher now performs the configured number of retries in synchronous runs. Previously, retry_attempts counted the initial request as an attempt, so synchronous runs performed one fewer retry than configured and behaved differently from asynchronous runs.LLMDocumentContentExtractor incorrectly treating a chat generator output that carries its own error field as a failed LLM call. Such documents are now processed normally.LLMMessagesRouter now raises a ValueError when output_names contains "chat_generator_text" or "unmatched". Those two outputs are always created by the router, so reusing either name used to be accepted and then silently broke the component: "chat_generator_text" had its socket retyped to list[ChatMessage] and lost the LLM reply at run time, while "unmatched" made a matched decision indistinguishable from an unmatched one.LLMMetadataExtractor incorrectly treating an extracted error metadata field as an extraction failure. error can now be used in expected_keys like any other metadata key.LLMRanker no longer crashes with AttributeError when query is None (or another non-string) coming from an upstream pipeline component. Both run and run_async now return the documents unranked, matching the existing empty-query fallback.MarkdownHeaderSplitter silently dropping a document's leading header line when secondary_split is used with keep_headers=False. This happened for any chunk that wasn't the result of an actual header split, for example when a document has no headers, only headers without body content, a header at a level excluded from header_split_levels, or input metadata that already contained a header key. The header line is now only stripped from chunks that truly came from a header split.MarkdownHeaderSplitter no longer drops the text that precedes the first header it splits on. A document's opening paragraph, title block, or front matter was silently lost, and so was any content above the first matching header when header_split_levels excluded the levels above it. A document whose headers all had empty bodies lost that text too, because it was treated as header-only. That text is now emitted as a leading chunk with empty header and parent_headers metadata, so every chunk still carries the documented metadata fields. Leading text that is only whitespace joins the first chunk instead of becoming a chunk of its own, which keeps the chunks a byte-exact partition of the input without adding an empty chunk.page_number metadata from MarkdownHeaderSplitter when secondary splitting uses overlap. Page breaks in overlapping content are no longer counted multiple times, and custom page_break_character values are respected.EmbeddingBasedDocumentSplitter emitting a final split shorter than min_length. Small splits were only merged forward, so the last one had nothing left to absorb and was returned as its own document. It is now merged into the preceding split, unless doing so would reach max_length, the same limit that already governs forward merges.MultiQueryTextRetriever, MultiQueryEmbeddingRetriever and MultiRetriever now honour max_workers in their run_async methods. Previously only the synchronous run bounded fan-out (via a ThreadPoolExecutor); run_async launched every sub-query / sub-retriever call at once, ignoring max_workers and risking rate limits or connection exhaustion against the underlying retriever or embedder. run_async now bounds concurrency with an asyncio.Semaphore, matching the behaviour of the LLM extractors.ChatMessage.from_openai_dict_format accepts the same empty content, so such a message round-trips.OpenAIDocumentEmbedder.run_async (and AzureOpenAIDocumentEmbedder) now request encoding_format="float" from the embeddings endpoint, matching the synchronous run. The async batch path was missed when this was added in #9655, so run_async let the OpenAI SDK negotiate the base64 wire format, which some OpenAI-compatible endpoints do not support.OpenAIChatGenerator and OpenAIResponsesChatGenerator so that an empty runtime tools list overrides tools configured at initialization.timeout and max_retries to the to_dict method of OpenAIImageGenerator. They were dropped on serialization, so a pipeline saved and reloaded silently fell back to the OPENAI_TIMEOUT/OPENAI_MAX_RETRIES defaults instead of the configured values. This matches OpenAIChatGenerator, OpenAITextEmbedder and OpenAIDocumentEmbedder, which already serialize both.OutputAdapter and ConditionalRouter no longer coerce string outputs to other Python types when output_type=str. Previously the rendered template was always passed through ast.literal_eval in safe mode, so a string that happened to be a valid Python literal was silently converted (for example "1,000" became the tuple (1, 0) and "42" became the integer 42), producing a value that did not match the declared str output type. Literal evaluation is now skipped when output_type is str, so string outputs are returned unchanged. Other output types still reconstruct structured literals as before.PythonCodeSplitter now deep-copies the metadata of the document it splits, matching the other splitters. Previously it copied it shallowly, so nested values such as a list under meta["tags"] were shared between every chunk and with the input document, and editing one chunk's metadata changed all the others. The secondary line-based split of oversized units had the same problem.RecursiveDocumentSplitter and TextCleaner now keep their init parameters when serialized, for example in a pipeline saved with Pipeline.dumps(). Before, RecursiveDocumentSplitter wrote split_unit: word whatever unit it was created with, so a reloaded splitter configured for characters or tokens counted words instead. A reloaded TextCleaner lost all of its options and returned the texts unchanged.DocumentSplitter with split_by="word", "period", "page", "passage", "line" or "sentence" generating redundant trailing chunks containing only overlap. Segments that add no new text beyond the already covered units are now skipped, mirroring the split_by="token" fix.IndexError: list index out of range when a chat generator's stream completes without emitting any chunks. _convert_streaming_chunks_to_chat_message now returns an assistant message with empty content and None metadata instead.SuperComponent.run_async now returns outputs mapped from components that are not leaves of the wrapped pipeline, like run does. Before, an output that was also consumed inside the pipeline, such as retriever.documents feeding a prompt builder, was silently missing from the async result. This also affected a PipelineTool run by Agent.run_async, whose outputs_to_state got nothing for such outputs.create_tool_from_function and the @tool decorator now work with functions defined in a module that uses from __future__ import annotations. Before, the parameter annotations were plain strings, and creating the tool raised a SchemaGenerationError for any parameter typed with something other than a builtin, such as Annotated, Literal, Optional, Document or State. The annotations are now resolved first, and Annotated descriptions are kept.MetadataRouter now raises a clear ValueError when a rule uses the reserved output name unmatched, instead of failing with an internal duplicate-keyword TypeError during output registration.FileTypeRouter and DocumentTypeRouter now raise a clear ValueError when a MIME type uses a reserved output name (unclassified, or failed on FileTypeRouter), instead of failing with an internal duplicate-keyword TypeError during output registration.XLSXToDocument writing the string nan into an empty cell when table_format="markdown". The same cell is written as an empty field by table_format="csv", so an empty cell now reads as empty in both formats. A different placeholder can be set with table_format_kwargs={"missingval": "N/A"}.@abhati27, @abo3losh1, @AjayShivran, @alanhuangyoo, @anakin87, @Anurag-M1, @Arman-Beykmohammadi, @ArzelaAscoIi, @Awshesh12, @bilgeyucel, @bogdankostic, @businessarshgoyal, @coder058, @CoralGarden52, @davidsbatista, @dfedoryshchev, @dingpuyu, @Diwak4r, @Dxfory, @eatLaoJun, @ege-arhan, @feiiiiii5, @fng713, @Goodnight77, @gyanu2507, @Harsh23Kashyap, @iridescentWen, @JbravoI, @JimmyWang0417, @jliounis, @julian-risch, @kacperlukawski, @Koushik890, @Kuang-xianxin, @L4XB, @Lesereingrape, @lets-order-some-fries, @linhongyu510, @mfurkanakinci, @mikemikimike, @MohammadHijjawi97, @mrchtr, @Nikhi00718, @otiscuilei, @PattonBrown, @pcbeingused333, @PeterSmith0127-lcm, @Rainmemery, @rautaditya2606, @Ricky-7-Yan, @sainikhiljuluri, @samrusani, @seanxuu, @ShousenZHANG, @simpleqt, @sjrl, @spacesheepinternet, @teachershuang, @tstadel, @Vedant-Agarwal, @vercel[bot], @winter-street, @xblwh, @yavuz-yilmaz
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
The OpenAI Chat Completions and Responses converters no longer raise on an assistant message with no content parts, which a Chat Generator returns whe
ChatMessage.from_openai_dict_format accepts the same empty content, so such a message round-trips.The OpenAI Chat Completions and Responses converters no longer raise on an assistant message with no content parts, which a Chat Generator returns whe
ChatMessage.from_openai_dict_format accepts the same empty content, so such a message round-trips.Fixed a remote code execution vulnerability that could be triggered by loading an untrusted pipeline in default safe mode ( Pipeline.load / Pipeline.l…
Experimental context compaction for Agent. Runs before LLM calls and shortens the conversation once it crosses a configured fraction of the context window. The built-in SlidingWindowCompactor first removes entire historical turns, then, if needed, individual steps of the current task, replacing what it removes with a short omission note.
A second experimental compaction strategy, ToolResultPruningCompactor, instead of dropping whole turns, replaces older/large tool results with placeholders while keeping the most recent tool-calling steps (and their parallel results) intact.
from haystack.hooks.compaction import CompactionHook, SlidingWindowCompactor, ToolResultPruningCompactor
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
hook = CompactionHook(
compactor=SlidingWindowCompactor(), # or ToolResultPruningCompactor(min_keep_steps=2, min_tokens=200),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-nano"),
tools=[web_search],
hooks={"before_llm": [hook]},
)New haystack.token_counters module for estimating message/tool-schema size before sending a request (e.g. to decide when to compact). ApproximateTokenCounter needs no dependencies; TiktokenCounter is a local counter, closer estimates for OpenAI models; OpenAITokenCounter calls OpenAI's counting API for exact, model-specific counts.
from haystack.token_counters import OpenAITokenCounter, ApproximateTokenCounter, TiktokenCounter
from haystack.dataclasses import ChatMessage
counter = OpenAITokenCounter("gpt-5-mini")
# No dependencies: estimates from text length.
counter_app = ApproximateTokenCounter(chars_per_token=4.0).count(messages)
# Closer for OpenAI models; needs: pip install tiktoken
counter_tiktoken = TiktokenCounter(encoding="o200k_base").count(messages)
count = counter.count([ChatMessage.from_user("Hello!")])AgentTool wraps a Haystack Agent as a Tool so another Agent can delegate to it. Only the wrapped Agent's final reply is visible to the caller; its intermediate steps remain outside the main agent's context.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import AgentTool
researcher = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-mini"),
system_prompt="You are a research specialist. Investigate the task and report your findings.",
tools=[web_search],
)
research_specialist = AgentTool(
agent=researcher,
name="research",
description="Research a question on the web and report the findings",
)
coordinator = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4"),
tools=[research_specialist],
system_prompt="You coordinate specialists. Delegate research questions, then answer the user.",
)
result = coordinator.run([ChatMessage.from_user("What are the latest developments in the Haystack framework?")])
print(result["last_message"].text)See the multi-agent systems tutorial for a full walkthrough.
Add an Agent.clone() method that returns a new Agent with the same configuration, optionally replacing some init parameters: variant = agent.clone(system_prompt="Answer in German.").
Added a link_format parameter to both PyPDFToDocument and PDFMinerToDocument components, matching the existing functionality in DOCXToDocument. Links are parsed from PDF annotations and appended at the bottom of the page content.
Agent now returns an exit_reason output reporting why the run stopped, making it easier to route the Agent's output downstream (for example with a ConditionalRouter). It is one of: "text" (the model returned a reply with no tool calls), the name of the tool that satisfied a tool exit condition (in which case last_message is that tool's result), or "max_agent_steps" (the Agent hit max_agent_steps before meeting an exit condition). The reason is also available to hooks via state.get("exit_reason"), so an after_run hook can, for instance, append a fallback answer when the step budget is exhausted.
Haystack components that use a document store now provide close and close_async methods for releasing resources. These methods are available on: AutoMergingRetriever, CacheChecker, DocumentWriter, FilterRetriever, and SentenceWindowRetriever. If the underlying Document Store does not implement the corresponding method, calling close or close_async has no effect.
Add a content-free haystack.agent.hook tracing span for every Agent hook invocation. Each span identifies the hook point, hook name, and hook type, allowing hook latency and failures to be attributed without tracing the potentially large Agent State. CompactionHook also adds its configured compaction strategy, estimated context size, whether compaction was triggered, its token target, and whether the compactor returned a replacement.
Added the HAYSTACK_UNSAFE_DESERIALIZATION environment variable as a process-wide equivalent of loading with unsafe=True. When set to a truthy value (1 or true), every Pipeline.load / Pipeline.loads / Pipeline.from_dict call skips all deserialization safety checks — the module allowlist, the builtin/import-primitive and control-plane denylists, the object-internals traversal guard, and the refusal to honor a component's own unsafe: true flag. This is intended for deployments that only ever load fully trusted pipelines and cannot pass unsafe=True at every call site. A warning is logged the first time it takes effect. Only enable it when every pipeline the process loads is trusted: a single untrusted pipeline then leads to arbitrary code execution.
The switch is not limited to pipeline loading: it disables the same checks for every deserialization path in the process, including ones that take no unsafe argument of their own — Tool.from_dict, State.from_dict (agent snapshot resume), and the ConditionalRouter and OutputAdapter Jinja sandbox flags. A deployment that loads trusted pipelines at startup but accepts serialized tools or agent state at request time is therefore exposed on the request path as well, not only at load time.
The variable is read once, on the first deserialization in the process, and the result is then frozen for the process lifetime: writes to it afterwards are ignored, in either direction. Set it before the first pipeline is loaded. Freezing keeps the safety mode of a process from changing under a caller's feet and, because the first read happens before any deserialized data can run, stops a hostile pipeline from switching the checks off while it is being loaded.
Serialized OutputAdapter and ConditionalRouter components containing Jinja custom_filters must now be loaded with Pipeline.load(..., unsafe=True) (or the equivalent Pipeline.loads / Pipeline.from_dict option).
exit_reason is now a reserved state key on Agent. If you defined a custom state_schema key named exit_reason, rename it: the Agent now raises a ValueError at initialization when a reserved key is redefined.
Agent.state_schema now contains the user-provided state schema, exactly as passed to __init__. Previously it contained the resolved schema, which also included messages and the keys the Agent manages internally (step_count, token_usage, exit_reason, ...). If you were reading agent.state_schema to inspect the effective runtime schema, use the new public attribute agent.resolved_state_schema instead.
DocumentMAPEvaluator scores can change because average precision now uses all unique, valid ground-truth comparison values as its denominator and credits each value at most once. Re-baseline evaluations that relied on the previous scores.
PipelineSnapshot.pipeline_state.inputs changed shape. It used to store one flattened value per socket, {component: {socket: value}}. It now stores the pipeline's internal inputs, keeping the component that sent each one in the order it arrived: {component: {socket: [{"sender": ..., "value": ...}]}}. PipelineState gained an inputs_format field recording which of the two shapes a snapshot uses. The same applies to BreakpointException.inputs, which returns that field.
You are affected if you read pipeline_state.inputs (or BreakpointException.inputs) directly, for example to display or post-process a snapshot. Resuming a pipeline with Pipeline.run(pipeline_snapshot=...) is not affected, and neither is code that only passes snapshots around or persists them.
To adapt, read the input from the list and take its value:
inputs = snapshot.pipeline_state.inputs["serialized_data"]
# before
value = inputs["my_component"]["my_socket"]
# now
value = inputs["my_component"]["my_socket"][0]["value"]A socket that received inputs from several senders has one list item per input, each recording the sender that produced it.
Snapshots written by earlier versions of Haystack have inputs_format set to None and keep the flattened shape, so branch on that field if you need to handle both.
Loading a serialized OutputAdapter or ConditionalRouter whose unsafe init parameter is set to true now raises DeserializationError unless the pipeline is loaded in unsafe mode. Pipelines that legitimately rely on an unsafe OutputAdapter/ConditionalRouter embedded in serialized data must now load with Pipeline.load(..., unsafe=True) (or Pipeline.loads / Pipeline.from_dict with unsafe=True).
Passing window_size=0 to SentenceWindowRetriever.run or SentenceWindowRetriever.run_async now raises a ValueError instead of silently using the window_size set in the constructor. You are affected if you pass window_size=0 at runtime, either directly or from an upstream component in a pipeline. If you were relying on 0 to mean "use the value from the constructor", omit the argument (or pass None) instead:
InMemoryDocumentStore.get_metadata_field_unique_values (and its async counterpart)'s search_term parameter now matches against the metadata field's own value (case-insensitive substring) instead of the document's content. Callers relying on the previous content-matching behavior will need to filter documents by content themselves before calling this method.
The Agent now warms up its hooks before every run, not only the first one, as it already does for Tools and Toolsets. If your hook has a warm_up() that does expensive setup (opening a client, loading a model), make it return early once done, for example if self._client is not None: return.
Haystack can call warm_up() on Tools and Toolsets more than once, for example before every run. Previously Toolset absorbed repeated calls with an internal _is_warmed_up flag; that flag is gone and every call now reaches your warm_up(). If your custom Tool or Toolset does expensive work there (connecting to a server, loading a model), or relied on the _is_warmed_up attribute, guard with your own state and return early, for example if self._client is not None: return.
Fixed the serialization of PDFMinerToDocument. The component did not define to_dict, so the default serialization fell back to reading the init parameters from same-named attributes. Since the layout parameters are stored in self.layout_params, they were silently serialized with their default values, for example a component created with char_margin=0.5 was serialized with char_margin=2.0. Custom layout parameters are now preserved when a pipeline is serialized and loaded again.
Cancel and await sibling retrieval tasks when a concurrent call fails in MultiRetriever, MultiQueryTextRetriever, or MultiQueryEmbeddingRetriever.
Fixed an infinite recursion in CSVDocumentSplitter when nested row and column blocks were split together.
Comparing two Document objects with == now takes all metadata into account. Previously, two documents with different metadata could be considered equal if the metadata contained keys with the same names as document fields (such as id or content).
Document.from_dict(document.to_dict()) now correctly rebuilds any document. Previously, if the metadata contained keys with the same names as document fields (such as id or meta), this either raised an error or silently lost those metadata entries.
Fixed DocumentNDCGEvaluator producing NDCG scores outside the documented 0.0 to 1.0 range when the same document appeared more than once. A document retrieved multiple times used to be counted multiple times, pushing the score above 1.0; a ground truth document listed multiple times used to inflate the ideal gain, keeping a perfect retrieval below 1.0. Each distinct relevant document is now counted once, with the same relevance, in both the actual and ideal gain, so scores stay within range.
Fix FilterPolicy.MERGE in InMemoryBM25Retriever and InMemoryEmbeddingRetriever so initialization filters are combined with runtime comparison filters instead of being silently overwritten.
Fixed AnswerBuilder returning referenced documents in a scrambled order instead of ascending source-index order. The referenced document indices were collected in a set and iterated directly, so documents were emitted in the set's internal hash-table order (e.g. citations [3] [10] [50] yielded documents ordered 10, 3, 50). This order was deterministic but did not match the intuitive source order. The referenced documents are now returned sorted by their source index.
Fixed serialize_type and deserialize_type to correctly round-trip Callable types that declare an explicit parameter list, such as Callable[[int, str], bool]. Previously the parameter list was dropped during serialization (producing a malformed string like typing.Callable[, bool]) and could no longer be deserialized. This affected components that serialize type annotations, for example ConditionalRouter and OutputAdapter using a Callable output type.
Fix DocumentMAPEvaluator to include missed relevant documents in the average precision denominator and avoid crediting duplicate retrievals of the same document.
Fixed DocumentSplitter producing chunks that were not present in the source document when split_threshold was set together with split_overlap. Merging a below-threshold trailing segment into the previous split re-appended the overlapping units, duplicating text. The overlap is now added only once.
Fixed EmbeddingBasedDocumentSplitter.run_async embedding through the synchronous path while recursively splitting chunks longer than max_length. Only the first pass was async: the recursion called the sync splitting helper, so the embedder's blocking run ran on the event loop for every over-long chunk. The recursion now embeds through run_async as well.
Fixed JSONConverter raising a KeyError instead of logging its intended "Failed to extract text, skipping it" warning when a source is a ByteStream without a file_path in its meta (for example ByteStream.from_string(...), the exact usage shown in the component's own docstring examples). Affected error paths: invalid UTF-8 content, a jq_schema filter that fails to apply, and malformed JSON content.
Fixed LinkContentFetcher rotating the User-Agent on a cursor shared by every URL in the same run()/run_async() call. The URLs are fetched concurrently, so a retry triggered by one of them advanced the user agent for the others, and each completed fetch reset the cursor for the requests still in flight — most retries went out with the un-rotated user agent. Each fetch now walks the user_agents list on its own, so a URL rotates exactly as documented no matter how many other URLs are fetched alongside it.
Fixed MarkdownHeaderSplitter silently dropping a trailing header that has no body text. With keep_headers=True (the default), a header at the end of the document whose only content is whitespace was buffered to prepend to the next chunk, but with no following chunk it was never emitted, so the split documents no longer reconstructed the original text. Such trailing headers are now emitted as a final chunk.
Fixed MarkdownHeaderSplitter collapsing blank lines that follow a header with no body text. With keep_headers=True, such headers were re-joined with a single newline when prepended to the next chunk, so the split documents did not reconstruct the original text. Chunk content is now sliced from the original text and is byte-exact.
Fixed MarkdownHeaderSplitter including surrounding whitespace in the header and parent_headers metadata fields. The header text is now stripped; chunk content still keeps the header line's original whitespace.
Fixed schema-based serialization of lists, tuples and sets holding mixed types. Previously the schema was derived from the first element only, so deserializing such a value raised an AttributeError or silently returned mis-typed data (for example an Agent State field or a pipeline breakpoint input holding [Document(...), "text", 3]). Mixed-type arrays now record one schema per position using the JSON Schema prefixItems keyword and round-trip correctly. Homogeneous arrays keep the exact same output as before, so existing snapshots still load.
Fixed MSGToDocument raising a KeyError when converting a ByteStream source that has no file_path in its meta (for example a bare ByteStream(data=...), rather than a file path or a stream produced via ByteStream.from_file_path). Attachments extracted from such a source no longer include a parent_file_path key, since there is no source file path to record.
Pipeline connections now always convert values in the same way. When a component output is connected to an input that accepts multiple types, Haystack is sometimes able to automatically convert the value, and more than one conversion may be possible. For example, a ChatMessage with text "hello" connected to an input annotated str | list[str] can be delivered either as plain text ("hello") or as a list containing the text (["hello"]). Previously the conversion strategy was chosen non-deterministically, so the same pipeline could return a different value across runs. The conversion strategy is now selected using a fixed priority: first, wrapping a value in a list or unwrapping a single-element list; second, converting between ChatMessage and str; and last, combining both conversions.
Fixed normalize_metadata (used by all file converters) returning the same dictionary object for every source when meta is None or a single dictionary. Each source now receives an independent copy, so mutating one source's metadata downstream no longer leaks into the others.
OpenAIResponsesChatGenerator no longer mutates the parameters schema of the Tool objects passed to it. Previously every run wrote additionalProperties: False into the user's live Tool.parameters, silently altering the tool for any other generator that shared the same Tool instance and making serialization round trips unstable.
OpenAIResponsesChatGenerator no longer raises IndexError when it is warmed up with an empty tools list.
Fixed the parent of the haystack.agent.step.tool spans when an Agent step invokes several tools. The parent span is now resolved once before the tools run, so all tool calls of a step appear as siblings. Previously each span asked the tracer for the currently active span from inside its own concurrent invocation, which made the tool calls after the first one appear nested under a sibling tool call.
Fixed the haystack.pipeline.output_data tracing tag being empty. The tag was set at the start of Pipeline.run/run_async from the still-empty outputs, and since tracing backends coerce a tag value when it is set, the recorded output was always an empty dictionary. It is now set once the run completes so it reflects the final pipeline outputs. The tag is also gated behind content tracing (HAYSTACK_CONTENT_TRACING_ENABLED), consistent with the component-level input/output tags.
Fixed resuming a Pipeline from a pipeline_snapshot that was taken on a component's second or later visit, which failed with PipelineComponentsBlockedError: Cannot run pipeline - all components are blocked. A snapshot stored only the values of the pipeline's inputs and dropped the information about which component had sent each one, so on resume every restored input looked like it came from outside the pipeline, and such an input can only trigger a component on its first visit. Snapshots now record the sender of each input. Snapshots created by earlier versions of Haystack behave as before, so re-create them to resume anywhere in a looping pipeline.
Fixed a resumed Pipeline passing malformed inputs to the component the snapshot was taken on, whenever that component ran more than once after the resume, for example inside a loop. Every visit after the first reused the handling meant only for the paused visit and skipped the regular input consumption, so a variadic component could receive a bare value where it expected a list, raising errors such as TypeError: object of type 'int' has no len() from a BranchJoiner. This affected snapshots taken at any visit count, including the first.
Fixed QueryExpander returning duplicate queries when the chat generator repeats an expansion. Generated queries are now deduplicated while preserving first-seen order, so repeated expansions no longer trigger redundant retrievals or consume the requested expansion budget. Both run and run_async are affected.
RecursiveDocumentSplitter's word-mode fixed-size fallback no longer counts a run of whitespace (e.g. a double space, tab, or page break) as a word, so it no longer produces chunks smaller than split_length. It also no longer emits a whitespace-only chunk when the text ends in whitespace right after a chunk boundary; that trailing whitespace is now attached to the previous chunk instead. This changes the exact chunk boundaries and chunk count produced by the word-unit fallback for any text containing such whitespace runs. Documents already split and indexed under the old behavior will produce different chunks if re-split after upgrading, so re-index any document store that relies on stable chunk boundaries from this fallback path.
Fixed an issue where PipelineBase.remove_component did not reset auto-variadic socket flags (is_lazy_variadic and wrap_input_in_list) on input sockets when components or connections were removed.
Fixed Pipeline.remove_component leaving dangling references to the removed component on the sockets of its neighboring components. Previously, removing a component reset only its own sockets, so a surviving neighbor kept the removed component's name in its input socket's senders (or output socket's receivers). This corrupted introspection and validation: Pipeline.inputs() hid a now-unconnected mandatory input, and feeding that input directly could raise a spurious "already connected" error. The removed component's name is now stripped from its neighbors' sockets as well.
SentenceWindowRetriever.run and SentenceWindowRetriever.run_async now validate an explicitly provided window_size=0 instead of treating it as unset and falling back to the constructor value.
Fixed the schema-aware serialization helper used for pipeline snapshots and Agent State (_serialize_value_with_schema) so it no longer silently passes unsupported objects through as if they were serialized. Values such as datetime, bytes, complex and arbitrary objects without a to_dict method were previously stored unchanged and mislabeled as strings, which broke JSON storage and round-tripping of snapshots. Unsupported values now raise a SerializationError, and the callers that build snapshots (pipeline breakpoints and State.to_dict) catch it to omit only the offending field while keeping the rest of the payload resumable.
Added support for serializing and deserializing frozenset values in _serialize_value_with_schema. A frozenset now round-trips back to a frozenset instead of being dropped.
Fixed serialize_type/deserialize_type for typing.Literal. Previously a Literal type hint was serialized with its values rendered as bare tokens (e.g. typing.Literal[yes, no]), which failed to deserialize, and values that looked like type names (e.g. Literal["int", "str"]) were silently turned into types on the round-trip. The values are now serialized with repr() and read back with ast.literal_eval, so a Literal type used by a component (such as OutputAdapter or ConditionalRouter) round-trips correctly through pipeline serialization.
Fixed an AttributeError: 'str' object has no attribute 'items' raised by create_tool_from_function, the @tool decorator, and ComponentTool when a tool parameter is named properties. Keys inside a JSON schema properties mapping are property names and are no longer misinterpreted as schema keywords when stripping the auto-generated title keywords.
Fixed create_tool_from_function, the @tool decorator, and ComponentTool corrupting a tool's JSON schema when the string title appears as a name rather than as a schema keyword. Stripping the auto-generated title keywords no longer deletes entries of $defs, definitions, patternProperties, dependentSchemas or dependentRequired (which would leave a $ref dangling or silently drop a validation rule), and no longer edits title keys inside default, const, enum or examples values, which are instance data and part of the tool's contract.
Fixed _ToolsetWrapper.__getitem__ (used when combining Toolsets with +) raising IndexError for negative indices, unlike a plain Toolset. Indexing a combined toolset now behaves consistently with a list of Tools, as documented.
Fixed serialization of types that contain ... (Ellipsis), such as variadic tuples (tuple[int, ...]) and Callable[..., X]. Previously serialize_type rendered the ... as the literal string "Ellipsis", which deserialize_type then rejected as a non-type builtin, so a component using such a type (for example OutputAdapter(output_type=tuple[int, ...])) could be serialized but not deserialized, breaking Pipeline.loads() / Pipeline.load(). These types now round-trip correctly, and pipelines serialized by older versions (which emitted "Ellipsis") can still be loaded.
Fixed ConfirmationHook applying a Human-in-the-Loop decision to the wrong tool call when a custom ConfirmationStrategy returns a decision with a missing or incorrect tool_call_id. Each decision is now bound to the tool call for which its strategy ran, and ID-bearing decisions are no longer matched to a different call by name. Haystack's existing requirement of exactly one decision per tool call is now explicitly enforced. Each matched decision is consumed after use, so a missing, unused, or reused decision raises a ValueError instead of being silently misapplied.
DocumentJoiner and AnswerJoiner now resolve top_k consistently and validate it. Previously, a runtime top_k=0 was treated as "unset" and silently fell back to the instance's top_k, instead of returning an empty list as requested. Both components now:
ValueError at initialization if top_k is not None and is less than or equal to 0.ValueError at runtime if top_k passed to run() is negative.run() is called with top_k=0, regardless of the instance's configured top_k.Fixes MetaFieldRanker silently treating a runtime top_k=0 as unset and falling back to the value configured at initialization. Runtime values that are not greater than zero now raise a ValueError as documented.
Fixed PythonCodeSplitter losing identifying context for oversized functions, methods, or classes. When a unit is too large and falls back to line-based secondary splitting, only the first resulting piece naturally retains the source def/class line; every piece now includes a qualified_name field in meta identifying the function, method, or class it came from.
Fixed RecursiveDocumentSplitter not setting the source_id meta field on the chunks it produces. It wrote only parent_id, while every other splitter in the library (DocumentSplitter, CSVDocumentSplitter, EmbeddingBasedDocumentSplitter, HierarchicalDocumentSplitter, MarkdownHeaderSplitter and PythonCodeSplitter) writes source_id. Components that follow that convention therefore rejected its output: SentenceWindowRetriever reads source_id by default and raises when it is absent, so it failed with "The retrieved documents must have 'source_id' in their metadata." on a pipeline that worked with any other splitter. Chunks now carry source_id as well as parent_id, which keeps its previous value for callers already reading it.
Tool functions defined in a module using from __future__ import annotations are now inspected correctly by Agent. Postponed annotations are stored as strings, so a parameter annotated with State was not recognized and the live State object was not injected into the tool call. The annotations are now resolved before they are inspected.
MarkdownHeaderSplitter and CSVDocumentSplitter now deep-copy the metadata of the document they split, matching DocumentSplitter. Previously they copied it shallowly, so nested values such as a list under meta["tags"] were shared between every chunk and with the input document, and editing one chunk's metadata changed all the others. HierarchicalDocumentSplitter had the same problem on its root node, which kept references into the input document's metadata.
Keep an Agent's execution counter in sync with step_count restored by a before_run hook, so restarted Agents continue from the saved step instead of resetting the count.
Prevent arbitrary callables registered as serialized Jinja custom filters from executing while a pipeline is loaded in safe mode. Jinja can invoke filters with constant arguments during template compilation, and its sandbox does not apply callable-safety checks to filters.
Harden callable deserialization by checking the real module of every object traversed in a dotted callable path. This prevents an allowlisted module from exposing an object defined in an unallowlisted module that leads back to an otherwise allowlisted final callable.
Fixed a remote code execution vulnerability that could be triggered by loading an untrusted pipeline in default safe mode (Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). A malicious pipeline could either (a) set unsafe: true on an OutputAdapter or ConditionalRouter to disable the Jinja sandbox entirely, or (b) register the thread_safe_import import primitive as a Jinja custom_filters entry to import os and execute arbitrary commands — bypassing the deserialization allowlist and Jinja sandbox. The fix denies import primitives during callable deserialization, refuses to honor a component's unsafe flag while loading in safe mode, and hardens the Jinja sandbox (OutputAdapter, ConditionalRouter, PromptBuilder, ChatPromptBuilder) to block attribute access on module objects and calls into dangerous modules.
Closed an additional remote code execution vector in the deserialization control-plane hardening: the pipeline loading entry points (Pipeline.loads / Pipeline.load / Pipeline.from_dict, which accept unsafe=True) and the execute primitives (Pipeline.run / run_async / run_async_generator / stream) were still resolvable from the allowlisted haystack namespace. Bound as a custom_filters entry on an OutputAdapter or ConditionalRouter (which bypass the Jinja sandbox), Pipeline.loads(..., unsafe=True) let a pipeline loaded in default safe mode load a nested pipeline whose own filters (allow_deserialization_module, deserialize_callable) bind under the nested unsafe context, disarming the process-wide allowlist with "*" and invoking os.system. All of these entry points are now marked as deserializer-internal, so they can never be produced by deserializing untrusted data. unsafe=True still bypasses the check by design; there are no public API changes.
Fixed a remote code execution vulnerability that could be triggered by loading an untrusted pipeline in default safe mode (Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). Because the deserialization allowlist admits the whole haystack namespace, the deserializer's own allowlist-administration function (allow_deserialization_module) and its resolution helpers (deserialize_callable, deserialize_type, import_class_by_name) were themselves resolvable from serialized data. A malicious pipeline could register allow_deserialization_module as a Jinja custom_filters entry (on an OutputAdapter or ConditionalRouter), call it with "*" to disarm the allowlist process-wide, and then use the equally-resolvable deserialize_callable to resolve and invoke os.system. Loading alone was enough to trigger this: a Jinja filter called with constant arguments runs while the component is being constructed, so the pipeline never had to be run. The same attribute walk could also reach the deserializer's mutable control-plane state directly — for example a filter bound to _extra_allowed_modules.append — to widen the allowlist persistently and stage a later attack. Relatedly, the handle resolver walked attribute names freely, so a handle could descend into object internals such as <function>.__globals__ (a live module namespace, and via it __builtins__ and eval/exec) or <type>.__subclasses__ — classic sandbox-escape gadgets that stay inside an allowlisted module. The fix refuses to deserialize the deserialization control plane as a whole: the allowlist administration and resolution helpers (marked at definition time), everything defined in haystack.core.serialization_security, and any bound method of the mutable allowlist/context state. It also refuses to traverse into dunder and frame/code attributes while resolving a handle. This applies to both the callable- and class-resolution paths, and is bypassed only when the pipeline is loaded with unsafe=True.
Harden FileSystemToolResultStore.read() so it only reads references that resolve within the configured store root. This closes a boundary gap where callers could previously pass an arbitrary filesystem path to read() instead of a store-scoped reference returned by write().
Extracted DOCXLinkFormat to a reusable LinkFormat Enum in haystack/components/converters/utils.py. DOCXLinkFormat is now an alias for backward compatibility.
The Agent now tracks an approximate current context-window size in its internal State under context_tokens, refreshed after every LLM call with that reply's prompt-plus-completion tokens (normalized across the prompt_tokens/completion_tokens and input_tokens/output_tokens key conventions). Unlike token_usage, which accumulates across the whole run, context_tokens is replaced each call. Hooks can read it via state.get("context_tokens") — for example, a before_llm hook that triggers context compaction once the value crosses a threshold. It is a best-effort snapshot: it is 0 when the generator does not report usage, and does not count messages appended after the latest call until the next call refreshes it.
before_run hooks can now read state.data["tools"]. The key was previously written only once the first step had started, so a before_run hook hit a KeyError. It holds a snapshot of the tools available at that point, refreshed before every LLM call, so with a dynamic toolset such as SearchableToolset a before_run hook sees only the tools discovered so far.
Added an opt-in strict_datetime_comparison keyword argument to document_matches_filter, InMemoryDocumentStore, and MetadataRouter. When enabled, timezone-naive and timezone-aware datetimes never match each other. By default, mixed-awareness datetimes continue to be reconciled by copying the timezone from the aware value to the naive one, and this behavior is now consistent across equality, membership, and ordering operators.
Added a filters parameter to InMemoryDocumentStore.get_metadata_field_unique_values (sync and async), allowing the set of documents considered when computing unique metadata field values to be restricted.
InMemoryDocumentStore.get_metadata_field_unique_values and its async counterpart now support pagination via from_ and size parameters, matching the behavior of other Document Stores (e.g. Chroma).
MockChatGenerator's response_fn can now be tool-aware. If the callable accepts a second positional argument, it also receives the tools passed to run/run_async (a ToolsType or None), so a dynamic mock can build tool calls whose arguments follow the tool's parameter schema or route between the available tools. Existing single-argument response_fn callables are unaffected and keep receiving only the messages.
State.to_dict now accepts a skip_keys parameter to exclude specific keys from the output.
Add a tracing span per ConfirmationStrategy run to ConfirmationHook. Each haystack.agent.hook.human_in_the_loop.strategy span identifies the tool call it confirms and records the strategy type and the applied confirm, modify, or reject decision. When content tracing is enabled, spans also carry the arguments the strategy was run with and the ToolExecutionDecision it returned under haystack.agent.hook.human_in_the_loop.strategy.input and haystack.agent.hook.human_in_the_loop.strategy.output; chat messages, the confirmation strategy context, and the Agent State are not recorded by the hook.
LLMEvaluator, LLMRanker, QueryExpander, LLMMetadataExtractor, LLMDocumentContentExtractor and LLMMessagesRouter now wrap their internal ChatGenerator calls in a haystack.chat_generator.run tracing span. These components do not return ChatMessage objects, so the LLM token usage carried in reply.meta["usage"] was previously lost to tracers. The new span exposes the generator's replies via the haystack.component.output tag, so token usage is now visible in traces (requires content tracing to be enabled). When the generator runs across threads, the span is nested under the component's span.
+ operator (toolset_a + toolset_b) and passing a Toolset to add() (toolset_a.add(toolset_b)). Pass Toolsets as a list wherever tools are accepted instead: Agent(tools=[toolset_a, toolset_b]).@Aarkin7, @anakin87, @anxkhn, @aquib8112, @Aryan-Pardeshi, @atikulmunna, @bharadwaj-pendyala, @bilgeyucel, @bogdankostic, @camgrimsec, @chuenchen309, @davidpavlovschi, @davidsbatista, @DhanushPillay, @DivyaNarahari97, @erikos, @GovindhKishore, @hxaxd, @immuhammadfurqan, @iridescentWen, @jaideeppyne, @julian-risch, @kacperlukawski, @KXHXK, @LHMQ878, @LK-maker-007, @lntutor, @manjunathbhaskar, @mittalpk, @MVS-source, @onatozmenn, @otiscuilei, @pcbeingused333, @rautaditya2606, @sjrl, @sohumt123, @Solaris-star, @TimurRakhmatullin86, @vidigoat, @vinkiYu, @winklemad, @yaodong-shen
Fixed a remote code execution vulnerability that could be triggered by loading an untrusted pipeline in default safe mode ( Pipeline.load / Pipeline.l…
OutputAdapter and ConditionalRouter components containing Jinja custom_filters must now be loaded with Pipeline.load(..., unsafe=True) (or the equivalent Pipeline.loads / Pipeline.from_dict option).Added the HAYSTACK_UNSAFE_DESERIALIZATION environment variable as a process-wide equivalent of loading with unsafe=True. When set to a truthy value (1 or true), every Pipeline.load / Pipeline.loads / Pipeline.from_dict call skips all deserialization safety checks — the module allowlist, the builtin/import-primitive and control-plane denylists, the object-internals traversal guard, and the refusal to honor a component's own unsafe: true flag. This is intended for deployments that only ever load fully trusted pipelines and cannot pass unsafe=True at every call site. A warning is logged the first time it takes effect. Only enable it when every pipeline the process loads is trusted: a single untrusted pipeline then leads to arbitrary code execution.
The switch is not limited to pipeline loading: it disables the same checks for every deserialization path in the process, including ones that take no unsafe argument of their own — Tool.from_dict, State.from_dict (agent snapshot resume), and the ConditionalRouter and OutputAdapter Jinja sandbox flags. A deployment that loads trusted pipelines at startup but accepts serialized tools or agent state at request time is therefore exposed on the request path as well, not only at load time.
The variable is read once, on the first deserialization in the process, and the result is then frozen for the process lifetime: writes to it afterwards are ignored, in either direction. Set it before the first pipeline is loaded. Freezing keeps the safety mode of a process from changing under a caller's feet and, because the first read happens before any deserialized data can run, stops a hostile pipeline from switching the checks off while it is being loaded.
exit_reason is now a reserved state key on Agent. If you defined a custom state_schema key named exit_reason, rename it: the Agent now raises a ValueError at initialization when a reserved key is redefined.
Agent.state_schema now contains the user-provided state schema, exactly as passed to __init__. Previously it contained the resolved schema, which also included messages and the keys the Agent manages internally (step_count, token_usage, exit_reason, ...). If you were reading agent.state_schema to inspect the effective runtime schema, use the new public attribute agent.resolved_state_schema instead.
DocumentMAPEvaluator scores can change because average precision now uses all unique, valid ground-truth comparison values as its denominator and credits each value at most once. Re-baseline evaluations that relied on the previous scores.
PipelineSnapshot.pipeline_state.inputs changed shape. It used to store one flattened value per socket, {component: {socket: value}}. It now stores the pipeline's internal inputs, keeping the component that sent each one in the order it arrived: {component: {socket: [{"sender": ..., "value": ...}]}}. PipelineState gained an inputs_format field recording which of the two shapes a snapshot uses. The same applies to BreakpointException.inputs, which returns that field.
You are affected if you read pipeline_state.inputs (or BreakpointException.inputs) directly, for example to display or post-process a snapshot. Resuming a pipeline with Pipeline.run(pipeline_snapshot=...) is not affected, and neither is code that only passes snapshots around or persists them.
To adapt, read the input from the list and take its value:
inputs = snapshot.pipeline_state.inputs["serialized_data"]
# before
value = inputs["my_component"]["my_socket"]
# now
value = inputs["my_component"]["my_socket"][0]["value"]A socket that received inputs from several senders has one list item per input, each recording the sender that produced it.
Snapshots written by earlier versions of Haystack have inputs_format set to None and keep the flattened shape, so branch on that field if you need to handle both.
Loading a serialized OutputAdapter or ConditionalRouter whose unsafe init parameter is set to true now raises DeserializationError unless the pipeline is loaded in unsafe mode. Pipelines that legitimately rely on an unsafe OutputAdapter/ConditionalRouter embedded in serialized data must now load with Pipeline.load(..., unsafe=True) (or Pipeline.loads / Pipeline.from_dict with unsafe=True).
Passing window_size=0 to SentenceWindowRetriever.run or SentenceWindowRetriever.run_async now raises a ValueError instead of silently using the window_size set in the constructor. You are affected if you pass window_size=0 at runtime, either directly or from an upstream component in a pipeline. If you were relying on 0 to mean "use the value from the constructor", omit the argument (or pass None) instead:
retriever = SentenceWindowRetriever(document_store=document_store, window_size=3)
# Before: silently used window_size=3
retriever.run(retrieved_documents=docs, window_size=0)
# After: omit the argument to use the constructor value
retriever.run(retrieved_documents=docs)InMemoryDocumentStore.get_metadata_field_unique_values (and its async counterpart)'s search_term parameter now matches against the metadata field's own value (case-insensitive substring) instead of the document's content. Callers relying on the previous content-matching behavior will need to filter documents by content themselves before calling this method.
The Agent now warms up its hooks before every run, not only the first one, as it already does for Tools and Toolsets. If your hook has a warm_up() that does expensive setup (opening a client, loading a model), make it return early once done, for example if self._client is not None: return.
Haystack can call warm_up() on Tools and Toolsets more than once, for example before every run. Previously Toolset absorbed repeated calls with an internal _is_warmed_up flag; that flag is gone and every call now reaches your warm_up(). If your custom Tool or Toolset does expensive work there (connecting to a server, loading a model), or relied on the _is_warmed_up attribute, guard with your own state and return early, for example if self._client is not None: return.
Added a link_format parameter to both PyPDFToDocument and PDFMinerToDocument components, matching the existing functionality in DOCXToDocument. Links are parsed from PDF annotations and appended at the bottom of the page content.
Added experimental context compaction for the Agent. CompactionHook runs before LLM calls and shortens the conversation when it reaches a configured fraction of the model's context window.
The first built-in strategy, SlidingWindowCompactor, preserves leading system messages, the latest user task, and as much complete recent conversation as the target allows. It removes complete historical turns first, and only when removing every historical turn is insufficient does it remove individual Agent steps from the current task. It replaces removed history with a short omission note, left where the removed messages used to sit: directly after the leading system messages when only historical turns were removed, and directly after the latest user message when the current task's own steps were removed. Only one note is ever present, because a later compaction folds an earlier one into its replacement.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.hooks.compaction import CompactionHook, SlidingWindowCompactor
hook = CompactionHook(
compactor=SlidingWindowCompactor(),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-nano"),
tools=[web_search],
hooks={"before_llm": [hook]},
)The hook uses provider-reported context usage when available and locally estimates the request otherwise, including tool schemas. Leave headroom above compact_at for the next reply and its tool results.
SlidingWindowCompactor treats an assistant message and its following tool results as one step, so a tool call is never separated from its results. Historical turns are likewise kept or removed in full, so an assistant reply is not retained without the user message it answers. It can also land above the requested target rather than under it, because leading system messages and the current task are never removed and min_keep_steps holds on to the newest Agent steps whatever their size, so a long system prompt or one large tool result can leave the conversation well over the target. Compaction is lossy: removed messages cannot be recovered or summarized by this strategy. Implement the Compactor protocol to provide a custom strategy.
CompactionHook and SlidingWindowCompactor emit an ExperimentalWarning and may change without a deprecation cycle.
Agent now returns an exit_reason output reporting why the run stopped, making it easier to route the Agent's output downstream (for example with a ConditionalRouter). It is one of: "text" (the model returned a reply with no tool calls), the name of the tool that satisfied a tool exit condition (in which case last_message is that tool's result), or "max_agent_steps" (the Agent hit max_agent_steps before meeting an exit condition). The reason is also available to hooks via state.get("exit_reason"), so an after_run hook can, for instance, append a fallback answer when the step budget is exhausted.
Add OpenAITokenCounter, which uses OpenAI's input token counting API to return model-specific counts for Haystack ChatMessage objects and optional tool schemas. Unlike local estimates, it supports OpenAI's exact accounting for request formatting, images, files, and tools.
Here is an example:
from haystack.dataclasses import ChatMessage
from haystack.token_counters import OpenAITokenCounter
counter = OpenAITokenCounter("gpt-5-mini")
count = counter.count([ChatMessage.from_user("Hello!")])Added haystack.token_counters: a TokenCounter protocol for estimating how many tokens a list of ChatMessage objects occupies, with two implementations.
Providers report token usage only after a call, and only for the call as a whole, so anything that needs a size beforehand - deciding whether a conversation still fits a model's context window, or how much of it to drop - has to estimate one.
from haystack.dataclasses import ChatMessage
from haystack.token_counters import ApproximateTokenCounter, TiktokenCounter
messages = [ChatMessage.from_user("Hello, how are you?")]
# No dependencies: estimates from text length.
ApproximateTokenCounter(chars_per_token=4.0).count(messages)
# Closer for OpenAI models; needs: pip install tiktoken
TiktokenCounter(encoding="o200k_base").count(messages)ApproximateTokenCounter needs nothing installed and estimates from text length at a configurable chars_per_token. TiktokenCounter counts with OpenAI's byte-pair encoder, which is closer for OpenAI models but requires tiktoken and drifts on other providers; it raises at construction when the dependency is missing, and loads its encoding on first use.
Neither can measure an image or a file, since a tokenizer only sees text and providers derive an image's cost from its dimensions. Both charge a flat tokens_per_image and tokens_per_file instead, counting images a tool returned inside its result as well as those a message carries directly. Raise those values if you send large images or long documents.
Tool schemas are sent alongside the messages and consume tokens too, so count takes an optional tools argument to have them included:
counter.count(messages, tools=[my_tool])Implement TokenCounter to count differently - for instance against a provider's own token-counting endpoint, which is the only way to have images counted exactly.
Added the experimental ToolResultPruningCompactor. It reduces Agent context usage by replacing older, large tool results with short placeholders while preserving tool-call/result structure. Results from a configurable number of recent tool-calling Agent steps remain intact, including parallel results from those steps.
from haystack.hooks.compaction import CompactionHook, ToolResultPruningCompactor
compaction_hook = CompactionHook(
compactor=ToolResultPruningCompactor(
min_keep_steps=2,
min_tokens=200,
),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)Add an Agent.clone() method that returns a new Agent with the same configuration, optionally replacing some init parameters: variant = agent.clone(system_prompt="Answer in German.").
Added AgentTool, a Tool that wraps a Haystack Agent, allowing it to be used as a tool by another Agent. It is a building block for multi-agent systems: an Agent specialized in one task becomes a tool that another Agent can delegate to. The calling Agent only sees the final reply, so all the steps the wrapped Agent takes stay out of its context. Sensible defaults make this work out of the box: the task is delegated as a single user message and comes back as text.
Example:
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import AgentTool, ComponentTool
from haystack_integrations.components.websearch.serperdev import SerperDevWebSearch
researcher = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-mini"),
system_prompt="You are a research specialist. Investigate the task and report your findings.",
tools=[
ComponentTool(
component=SerperDevWebSearch(
top_k=3,
),
name="web_search",
description="Search the web for current information on any topic",
),
],
)
research = AgentTool(
agent=researcher,
name="research",
description="Research a question on the web and report the findings",
)
coordinator = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4"),
tools=[research],
system_prompt="You coordinate specialists. Delegate research questions, then answer the user.",
)
result = coordinator.run([ChatMessage.from_user("What are the latest developments in the Haystack framework?")])
print(result["last_message"].text)Haystack components that use a document store now provide close and close_async methods for releasing resources. These methods are available on: AutoMergingRetriever, CacheChecker, DocumentWriter, FilterRetriever, and SentenceWindowRetriever. If the underlying Document Store does not implement the corresponding method, calling close or close_async has no effect.
Add a content-free haystack.agent.hook tracing span for every Agent hook invocation. Each span identifies the hook point, hook name, and hook type, allowing hook latency and failures to be attributed without tracing the potentially large Agent State. CompactionHook also adds its configured compaction strategy, estimated context size, whether compaction was triggered, its token target, and whether the compactor returned a replacement.
DOCXLinkFormat to a reusable LinkFormat Enum in haystack/components/converters/utils.py. DOCXLinkFormat is now an alias for backward compatibility.Agent now tracks an approximate current context-window size in its internal State under context_tokens, refreshed after every LLM call with that reply's prompt-plus-completion tokens (normalized across the prompt_tokens/completion_tokens and input_tokens/output_tokens key conventions). Unlike token_usage, which accumulates across the whole run, context_tokens is replaced each call. Hooks can read it via state.get("context_tokens") — for example, a before_llm hook that triggers context compaction once the value crosses a threshold. It is a best-effort snapshot: it is 0 when the generator does not report usage, and does not count messages appended after the latest call until the next call refreshes it.before_run hooks can now read state.data["tools"]. The key was previously written only once the first step had started, so a before_run hook hit a KeyError. It holds a snapshot of the tools available at that point, refreshed before every LLM call, so with a dynamic toolset such as SearchableToolset a before_run hook sees only the tools discovered so far.strict_datetime_comparison keyword argument to document_matches_filter, InMemoryDocumentStore, and MetadataRouter. When enabled, timezone-naive and timezone-aware datetimes never match each other. By default, mixed-awareness datetimes continue to be reconciled by copying the timezone from the aware value to the naive one, and this behavior is now consistent across equality, membership, and ordering operators.filters parameter to InMemoryDocumentStore.get_metadata_field_unique_values (sync and async), allowing the set of documents considered when computing unique metadata field values to be restricted.InMemoryDocumentStore.get_metadata_field_unique_values and its async counterpart now support pagination via from_ and size parameters, matching the behavior of other Document Stores (e.g. Chroma).MockChatGenerator's response_fn can now be tool-aware. If the callable accepts a second positional argument, it also receives the tools passed to run/run_async (a ToolsType or None), so a dynamic mock can build tool calls whose arguments follow the tool's parameter schema or route between the available tools. Existing single-argument response_fn callables are unaffected and keep receiving only the messages.State.to_dict now accepts a skip_keys parameter to exclude specific keys from the output.ConfirmationStrategy run to ConfirmationHook. Each haystack.agent.hook.human_in_the_loop.strategy span identifies the tool call it confirms and records the strategy type and the applied confirm, modify, or reject decision. When content tracing is enabled, spans also carry the arguments the strategy was run with and the ToolExecutionDecision it returned under haystack.agent.hook.human_in_the_loop.strategy.input and haystack.agent.hook.human_in_the_loop.strategy.output; chat messages, the confirmation strategy context, and the Agent State are not recorded by the hook.LLMEvaluator, LLMRanker, QueryExpander, LLMMetadataExtractor, LLMDocumentContentExtractor and LLMMessagesRouter now wrap their internal ChatGenerator calls in a haystack.chat_generator.run tracing span. These components do not return ChatMessage objects, so the LLM token usage carried in reply.meta["usage"] was previously lost to tracers. The new span exposes the generator's replies via the haystack.component.output tag, so token usage is now visible in traces (requires content tracing to be enabled). When the generator runs across threads, the span is nested under the component's span.+ operator (toolset_a + toolset_b) and passing a Toolset to add() (toolset_a.add(toolset_b)). Pass Toolsets as a list wherever tools are accepted instead: Agent(tools=[toolset_a, toolset_b]).Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). A malicious pipeline could either (a) set unsafe: true on an OutputAdapter or ConditionalRouter to disable the Jinja sandbox entirely, or (b) register the thread_safe_import import primitive as a Jinja custom_filters entry to import os and execute arbitrary commands — bypassing the deserialization allowlist and Jinja sandbox. The fix denies import primitives during callable deserialization, refuses to honor a component's unsafe flag while loading in safe mode, and hardens the Jinja sandbox (OutputAdapter, ConditionalRouter, PromptBuilder, ChatPromptBuilder) to block attribute access on module objects and calls into dangerous modules.Pipeline.loads / Pipeline.load / Pipeline.from_dict, which accept unsafe=True) and the execute primitives (Pipeline.run / run_async / run_async_generator / stream) were still resolvable from the allowlisted haystack namespace. Bound as a custom_filters entry on an OutputAdapter or ConditionalRouter (which bypass the Jinja sandbox), Pipeline.loads(..., unsafe=True) let a pipeline loaded in default safe mode load a nested pipeline whose own filters (allow_deserialization_module, deserialize_callable) bind under the nested unsafe context, disarming the process-wide allowlist with "*" and invoking os.system. All of these entry points are now marked as deserializer-internal, so they can never be produced by deserializing untrusted data. unsafe=True still bypasses the check by design; there are no public API changes.Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). Because the deserialization allowlist admits the whole haystack namespace, the deserializer's own allowlist-administration function (allow_deserialization_module) and its resolution helpers (deserialize_callable, deserialize_type, import_class_by_name) were themselves resolvable from serialized data. A malicious pipeline could register allow_deserialization_module as a Jinja custom_filters entry (on an OutputAdapter or ConditionalRouter), call it with "*" to disarm the allowlist process-wide, and then use the equally-resolvable deserialize_callable to resolve and invoke os.system. Loading alone was enough to trigger this: a Jinja filter called with constant arguments runs while the component is being constructed, so the pipeline never had to be run. The same attribute walk could also reach the deserializer's mutable control-plane state directly — for example a filter bound to _extra_allowed_modules.append — to widen the allowlist persistently and stage a later attack. Relatedly, the handle resolver walked attribute names freely, so a handle could descend into object internals such as <function>.__globals__ (a live module namespace, and via it __builtins__ and eval/exec) or <type>.__subclasses__ — classic sandbox-escape gadgets that stay inside an allowlisted module. The fix refuses to deserialize the deserialization control plane as a whole: the allowlist administration and resolution helpers (marked at definition time), everything defined in haystack.core.serialization_security, and any bound method of the mutable allowlist/context state. It also refuses to traverse into dunder and frame/code attributes while resolving a handle. This applies to both the callable- and class-resolution paths, and is bypassed only when the pipeline is loaded with unsafe=True.FileSystemToolResultStore.read() so it only reads references that resolve within the configured store root. This closes a boundary gap where callers could previously pass an arbitrary filesystem path to read() instead of a store-scoped reference returned by write().Fixed the serialization of PDFMinerToDocument. The component did not define to_dict, so the default serialization fell back to reading the init parameters from same-named attributes. Since the layout parameters are stored in self.layout_params, they were silently serialized with their default values, for example a component created with char_margin=0.5 was serialized with char_margin=2.0. Custom layout parameters are now preserved when a pipeline is serialized and loaded again.
Cancel and await sibling retrieval tasks when a concurrent call fails in MultiRetriever, MultiQueryTextRetriever, or MultiQueryEmbeddingRetriever.
Fixed an infinite recursion in CSVDocumentSplitter when nested row and column blocks were split together.
Comparing two Document objects with == now takes all metadata into account. Previously, two documents with different metadata could be considered equal if the metadata contained keys with the same names as document fields (such as id or content).
Document.from_dict(document.to_dict()) now correctly rebuilds any document. Previously, if the metadata contained keys with the same names as document fields (such as id or meta), this either raised an error or silently lost those metadata entries.
Fixed DocumentNDCGEvaluator producing NDCG scores outside the documented 0.0 to 1.0 range when the same document appeared more than once. A document retrieved multiple times used to be counted multiple times, pushing the score above 1.0; a ground truth document listed multiple times used to inflate the ideal gain, keeping a perfect retrieval below 1.0. Each distinct relevant document is now counted once, with the same relevance, in both the actual and ideal gain, so scores stay within range.
Fixed AnswerBuilder returning referenced documents in a scrambled order instead of ascending source-index order. The referenced document indices were collected in a set and iterated directly, so documents were emitted in the set's internal hash-table order (e.g. citations [3] [10] [50] yielded documents ordered 10, 3, 50). This order was deterministic but did not match the intuitive source order. The referenced documents are now returned sorted by their source index.
Fixed serialize_type and deserialize_type to correctly round-trip Callable types that declare an explicit parameter list, such as Callable[[int, str], bool]. Previously the parameter list was dropped during serialization (producing a malformed string like typing.Callable[, bool]) and could no longer be deserialized. This affected components that serialize type annotations, for example ConditionalRouter and OutputAdapter using a Callable output type.
Fix DocumentMAPEvaluator to include missed relevant documents in the average precision denominator and avoid crediting duplicate retrievals of the same document.
Fixed DocumentSplitter producing chunks that were not present in the source document when split_threshold was set together with split_overlap. Merging a below-threshold trailing segment into the previous split re-appended the overlapping units, duplicating text. The overlap is now added only once.
Fixed EmbeddingBasedDocumentSplitter.run_async embedding through the synchronous path while recursively splitting chunks longer than max_length. Only the first pass was async: the recursion called the sync splitting helper, so the embedder's blocking run ran on the event loop for every over-long chunk. The recursion now embeds through run_async as well.
Fixed JSONConverter raising a KeyError instead of logging its intended "Failed to extract text, skipping it" warning when a source is a ByteStream without a file_path in its meta (for example ByteStream.from_string(...), the exact usage shown in the component's own docstring examples). Affected error paths: invalid UTF-8 content, a jq_schema filter that fails to apply, and malformed JSON content.
Fixed LinkContentFetcher rotating the User-Agent on a cursor shared by every URL in the same run()/run_async() call. The URLs are fetched concurrently, so a retry triggered by one of them advanced the user agent for the others, and each completed fetch reset the cursor for the requests still in flight — most retries went out with the un-rotated user agent. Each fetch now walks the user_agents list on its own, so a URL rotates exactly as documented no matter how many other URLs are fetched alongside it.
Fixed MarkdownHeaderSplitter silently dropping a trailing header that has no body text. With keep_headers=True (the default), a header at the end of the document whose only content is whitespace was buffered to prepend to the next chunk, but with no following chunk it was never emitted, so the split documents no longer reconstructed the original text. Such trailing headers are now emitted as a final chunk.
Fixed MarkdownHeaderSplitter collapsing blank lines that follow a header with no body text. With keep_headers=True, such headers were re-joined with a single newline when prepended to the next chunk, so the split documents did not reconstruct the original text. Chunk content is now sliced from the original text and is byte-exact.
Fixed MarkdownHeaderSplitter including surrounding whitespace in the header and parent_headers metadata fields. The header text is now stripped; chunk content still keeps the header line's original whitespace.
Fixed schema-based serialization of lists, tuples and sets holding mixed types. Previously the schema was derived from the first element only, so deserializing such a value raised an AttributeError or silently returned mis-typed data (for example an Agent State field or a pipeline breakpoint input holding [Document(...), "text", 3]). Mixed-type arrays now record one schema per position using the JSON Schema prefixItems keyword and round-trip correctly. Homogeneous arrays keep the exact same output as before, so existing snapshots still load.
Fixed MSGToDocument raising a KeyError when converting a ByteStream source that has no file_path in its meta (for example a bare ByteStream(data=...), rather than a file path or a stream produced via ByteStream.from_file_path). Attachments extracted from such a source no longer include a parent_file_path key, since there is no source file path to record.
Pipeline connections now always convert values in the same way. When a component output is connected to an input that accepts multiple types, Haystack is sometimes able to automatically convert the value, and more than one conversion may be possible. For example, a ChatMessage with text "hello" connected to an input annotated str | list[str] can be delivered either as plain text ("hello") or as a list containing the text (["hello"]). Previously the conversion strategy was chosen non-deterministically, so the same pipeline could return a different value across runs. The conversion strategy is now selected using a fixed priority: first, wrapping a value in a list or unwrapping a single-element list; second, converting between ChatMessage and str; and last, combining both conversions.
Fixed normalize_metadata (used by all file converters) returning the same dictionary object for every source when meta is None or a single dictionary. Each source now receives an independent copy, so mutating one source's metadata downstream no longer leaks into the others.
OpenAIResponsesChatGenerator no longer mutates the parameters schema of the Tool objects passed to it. Previously every run wrote additionalProperties: False into the user's live Tool.parameters, silently altering the tool for any other generator that shared the same Tool instance and making serialization round trips unstable.
OpenAIResponsesChatGenerator no longer raises IndexError when it is warmed up with an empty tools list.
Fixed the parent of the haystack.agent.step.tool spans when an Agent step invokes several tools. The parent span is now resolved once before the tools run, so all tool calls of a step appear as siblings. Previously each span asked the tracer for the currently active span from inside its own concurrent invocation, which made the tool calls after the first one appear nested under a sibling tool call.
Fixed the haystack.pipeline.output_data tracing tag being empty. The tag was set at the start of Pipeline.run/run_async from the still-empty outputs, and since tracing backends coerce a tag value when it is set, the recorded output was always an empty dictionary. It is now set once the run completes so it reflects the final pipeline outputs. The tag is also gated behind content tracing (HAYSTACK_CONTENT_TRACING_ENABLED), consistent with the component-level input/output tags.
Fixed resuming a Pipeline from a pipeline_snapshot that was taken on a component's second or later visit, which failed with PipelineComponentsBlockedError: Cannot run pipeline - all components are blocked. A snapshot stored only the values of the pipeline's inputs and dropped the information about which component had sent each one, so on resume every restored input looked like it came from outside the pipeline, and such an input can only trigger a component on its first visit. Snapshots now record the sender of each input. Snapshots created by earlier versions of Haystack behave as before, so re-create them to resume anywhere in a looping pipeline.
Fixed a resumed Pipeline passing malformed inputs to the component the snapshot was taken on, whenever that component ran more than once after the resume, for example inside a loop. Every visit after the first reused the handling meant only for the paused visit and skipped the regular input consumption, so a variadic component could receive a bare value where it expected a list, raising errors such as TypeError: object of type 'int' has no len() from a BranchJoiner. This affected snapshots taken at any visit count, including the first.
Fixed QueryExpander returning duplicate queries when the chat generator repeats an expansion. Generated queries are now deduplicated while preserving first-seen order, so repeated expansions no longer trigger redundant retrievals or consume the requested expansion budget. Both run and run_async are affected.
RecursiveDocumentSplitter's word-mode fixed-size fallback no longer counts a run of whitespace (e.g. a double space, tab, or page break) as a word, so it no longer produces chunks smaller than split_length. It also no longer emits a whitespace-only chunk when the text ends in whitespace right after a chunk boundary; that trailing whitespace is now attached to the previous chunk instead.
This changes the exact chunk boundaries and chunk count produced by the word-unit fallback for any text containing such whitespace runs. Documents already split and indexed under the old behavior will produce different chunks if re-split after upgrading, so re-index any document store that relies on stable chunk boundaries from this fallback path.
Fixed an issue where PipelineBase.remove_component did not reset auto-variadic socket flags (is_lazy_variadic and wrap_input_in_list) on input sockets when components or connections were removed.
Fixed Pipeline.remove_component leaving dangling references to the removed component on the sockets of its neighboring components. Previously, removing a component reset only its own sockets, so a surviving neighbor kept the removed component's name in its input socket's senders (or output socket's receivers). This corrupted introspection and validation: Pipeline.inputs() hid a now-unconnected mandatory input, and feeding that input directly could raise a spurious "already connected" error. The removed component's name is now stripped from its neighbors' sockets as well.
SentenceWindowRetriever.run and SentenceWindowRetriever.run_async now validate an explicitly provided window_size=0 instead of treating it as unset and falling back to the constructor value.
Fixed the schema-aware serialization helper used for pipeline snapshots and Agent State (_serialize_value_with_schema) so it no longer silently passes unsupported objects through as if they were serialized. Values such as datetime, bytes, complex and arbitrary objects without a to_dict method were previously stored unchanged and mislabeled as strings, which broke JSON storage and round-tripping of snapshots. Unsupported values now raise a SerializationError, and the callers that build snapshots (pipeline breakpoints and State.to_dict) catch it to omit only the offending field while keeping the rest of the payload resumable.
Added support for serializing and deserializing frozenset values in _serialize_value_with_schema. A frozenset now round-trips back to a frozenset instead of being dropped.
Fixed serialize_type/deserialize_type for typing.Literal. Previously a Literal type hint was serialized with its values rendered as bare tokens (e.g. typing.Literal[yes, no]), which failed to deserialize, and values that looked like type names (e.g. Literal["int", "str"]) were silently turned into types on the round-trip. The values are now serialized with repr() and read back with ast.literal_eval, so a Literal type used by a component (such as OutputAdapter or ConditionalRouter) round-trips correctly through pipeline serialization.
Fixed an AttributeError: 'str' object has no attribute 'items' raised by create_tool_from_function, the @tool decorator, and ComponentTool when a tool parameter is named properties. Keys inside a JSON schema properties mapping are property names and are no longer misinterpreted as schema keywords when stripping the auto-generated title keywords.
Fixed create_tool_from_function, the @tool decorator, and ComponentTool corrupting a tool's JSON schema when the string title appears as a name rather than as a schema keyword. Stripping the auto-generated title keywords no longer deletes entries of $defs, definitions, patternProperties, dependentSchemas or dependentRequired (which would leave a $ref dangling or silently drop a validation rule), and no longer edits title keys inside default, const, enum or examples values, which are instance data and part of the tool's contract.
Fixed _ToolsetWrapper.__getitem__ (used when combining Toolsets with +) raising IndexError for negative indices, unlike a plain Toolset. Indexing a combined toolset now behaves consistently with a list of Tools, as documented.
Fixed serialization of types that contain ... (Ellipsis), such as variadic tuples (tuple[int, ...]) and Callable[..., X]. Previously serialize_type rendered the ... as the literal string "Ellipsis", which deserialize_type then rejected as a non-type builtin, so a component using such a type (for example OutputAdapter(output_type=tuple[int, ...])) could be serialized but not deserialized, breaking Pipeline.loads() / Pipeline.load(). These types now round-trip correctly, and pipelines serialized by older versions (which emitted "Ellipsis") can still be loaded.
Fixed ConfirmationHook applying a Human-in-the-Loop decision to the wrong tool call when a custom ConfirmationStrategy returns a decision with a missing or incorrect tool_call_id. Each decision is now bound to the tool call for which its strategy ran, and ID-bearing decisions are no longer matched to a different call by name. Haystack's existing requirement of exactly one decision per tool call is now explicitly enforced. Each matched decision is consumed after use, so a missing, unused, or reused decision raises a ValueError instead of being silently misapplied.
DocumentJoiner and AnswerJoiner now resolve top_k consistently and validate it. Previously, a runtime top_k=0 was treated as "unset" and silently fell back to the instance's top_k, instead of returning an empty list as requested. Both components now:
ValueError at initialization if top_k is not None and is less than or equal to 0.ValueError at runtime if top_k passed to run() is negative.run() is called with top_k=0, regardless of the instance's configured top_k.Fixes MetaFieldRanker silently treating a runtime top_k=0 as unset and falling back to the value configured at initialization. Runtime values that are not greater than zero now raise a ValueError as documented.
Fixed PythonCodeSplitter losing identifying context for oversized functions, methods, or classes. When a unit is too large and falls back to line-based secondary splitting, only the first resulting piece naturally retains the source def/class line; every piece now includes a qualified_name field in meta identifying the function, method, or class it came from.
Fixed RecursiveDocumentSplitter not setting the source_id meta field on the chunks it produces. It wrote only parent_id, while every other splitter in the library (DocumentSplitter, CSVDocumentSplitter, EmbeddingBasedDocumentSplitter, HierarchicalDocumentSplitter, MarkdownHeaderSplitter and PythonCodeSplitter) writes source_id. Components that follow that convention therefore rejected its output: SentenceWindowRetriever reads source_id by default and raises when it is absent, so it failed with "The retrieved documents must have 'source_id' in their metadata." on a pipeline that worked with any other splitter. Chunks now carry source_id as well as parent_id, which keeps its previous value for callers already reading it.
Tool functions defined in a module using from __future__ import annotations are now inspected correctly by Agent. Postponed annotations are stored as strings, so a parameter annotated with State was not recognized and the live State object was not injected into the tool call. The annotations are now resolved before they are inspected.
MarkdownHeaderSplitter and CSVDocumentSplitter now deep-copy the metadata of the document they split, matching DocumentSplitter. Previously they copied it shallowly, so nested values such as a list under meta["tags"] were shared between every chunk and with the input document, and editing one chunk's metadata changed all the others. HierarchicalDocumentSplitter had the same problem on its root node, which kept references into the input document's metadata.
Keep an Agent's execution counter in sync with step_count restored by a before_run hook, so restarted Agents continue from the saved step instead of resetting the count.
@Aarkin7, @anakin87, @anxkhn, @aquib8112, @Aryan-Pardeshi, @atikulmunna, @bharadwaj-pendyala, @bilgeyucel, @bogdankostic, @camgrimsec, @chuenchen309, @davidpavlovschi, @davidsbatista, @DhanushPillay, @DivyaNarahari97, @erikos, @GovindhKishore, @hxaxd, @immuhammadfurqan, @iridescentWen, @jaideeppyne, @julian-risch, @kacperlukawski, @KXHXK, @LHMQ878, @LK-maker-007, @lntutor, @manjunathbhaskar, @mittalpk, @MVS-source, @onatozmenn, @otiscuilei, @pcbeingused333, @rautaditya2606, @sjrl, @sohumt123, @Solaris-star, @TimurRakhmatullin86, @vidigoat, @vinkiYu, @winklemad, @yaodong-shen
Fixed a remote code execution vulnerability that could be triggered by loading an untrusted pipeline in default safe mode ( Pipeline.load / Pipeline.l…
OutputAdapter and ConditionalRouter components containing Jinja custom_filters must now be loaded with Pipeline.load(..., unsafe=True) (or the equivalent Pipeline.loads / Pipeline.from_dict option).Added the HAYSTACK_UNSAFE_DESERIALIZATION environment variable as a process-wide equivalent of loading with unsafe=True. When set to a truthy value (1 or true), every Pipeline.load / Pipeline.loads / Pipeline.from_dict call skips all deserialization safety checks — the module allowlist, the builtin/import-primitive and control-plane denylists, the object-internals traversal guard, and the refusal to honor a component's own unsafe: true flag. This is intended for deployments that only ever load fully trusted pipelines and cannot pass unsafe=True at every call site. A warning is logged the first time it takes effect. Only enable it when every pipeline the process loads is trusted: a single untrusted pipeline then leads to arbitrary code execution.
The switch is not limited to pipeline loading: it disables the same checks for every deserialization path in the process, including ones that take no unsafe argument of their own — Tool.from_dict, State.from_dict (agent snapshot resume), and the ConditionalRouter and OutputAdapter Jinja sandbox flags. A deployment that loads trusted pipelines at startup but accepts serialized tools or agent state at request time is therefore exposed on the request path as well, not only at load time.
The variable is read once, on the first deserialization in the process, and the result is then frozen for the process lifetime: writes to it afterwards are ignored, in either direction. Set it before the first pipeline is loaded. Freezing keeps the safety mode of a process from changing under a caller's feet and, because the first read happens before any deserialized data can run, stops a hostile pipeline from switching the checks off while it is being loaded.
exit_reason is now a reserved state key on Agent. If you defined a custom state_schema key named exit_reason, rename it: the Agent now raises a ValueError at initialization when a reserved key is redefined.
Agent.state_schema now contains the user-provided state schema, exactly as passed to __init__. Previously it contained the resolved schema, which also included messages and the keys the Agent manages internally (step_count, token_usage, exit_reason, ...). If you were reading agent.state_schema to inspect the effective runtime schema, use the new public attribute agent.resolved_state_schema instead.
DocumentMAPEvaluator scores can change because average precision now uses all unique, valid ground-truth comparison values as its denominator and credits each value at most once. Re-baseline evaluations that relied on the previous scores.
PipelineSnapshot.pipeline_state.inputs changed shape. It used to store one flattened value per socket, {component: {socket: value}}. It now stores the pipeline's internal inputs, keeping the component that sent each one in the order it arrived: {component: {socket: [{"sender": ..., "value": ...}]}}. PipelineState gained an inputs_format field recording which of the two shapes a snapshot uses. The same applies to BreakpointException.inputs, which returns that field.
You are affected if you read pipeline_state.inputs (or BreakpointException.inputs) directly, for example to display or post-process a snapshot. Resuming a pipeline with Pipeline.run(pipeline_snapshot=...) is not affected, and neither is code that only passes snapshots around or persists them.
To adapt, read the input from the list and take its value:
inputs = snapshot.pipeline_state.inputs["serialized_data"]
# before
value = inputs["my_component"]["my_socket"]
# now
value = inputs["my_component"]["my_socket"][0]["value"]A socket that received inputs from several senders has one list item per input, each recording the sender that produced it.
Snapshots written by earlier versions of Haystack have inputs_format set to None and keep the flattened shape, so branch on that field if you need to handle both.
Loading a serialized OutputAdapter or ConditionalRouter whose unsafe init parameter is set to true now raises DeserializationError unless the pipeline is loaded in unsafe mode. Pipelines that legitimately rely on an unsafe OutputAdapter/ConditionalRouter embedded in serialized data must now load with Pipeline.load(..., unsafe=True) (or Pipeline.loads / Pipeline.from_dict with unsafe=True).
Passing window_size=0 to SentenceWindowRetriever.run or SentenceWindowRetriever.run_async now raises a ValueError instead of silently using the window_size set in the constructor. You are affected if you pass window_size=0 at runtime, either directly or from an upstream component in a pipeline. If you were relying on 0 to mean "use the value from the constructor", omit the argument (or pass None) instead:
retriever = SentenceWindowRetriever(document_store=document_store, window_size=3)
# Before: silently used window_size=3
retriever.run(retrieved_documents=docs, window_size=0)
# After: omit the argument to use the constructor value
retriever.run(retrieved_documents=docs)InMemoryDocumentStore.get_metadata_field_unique_values (and its async counterpart)'s search_term parameter now matches against the metadata field's own value (case-insensitive substring) instead of the document's content. Callers relying on the previous content-matching behavior will need to filter documents by content themselves before calling this method.
The Agent now warms up its hooks before every run, not only the first one, as it already does for Tools and Toolsets. If your hook has a warm_up() that does expensive setup (opening a client, loading a model), make it return early once done, for example if self._client is not None: return.
Haystack can call warm_up() on Tools and Toolsets more than once, for example before every run. Previously Toolset absorbed repeated calls with an internal _is_warmed_up flag; that flag is gone and every call now reaches your warm_up(). If your custom Tool or Toolset does expensive work there (connecting to a server, loading a model), or relied on the _is_warmed_up attribute, guard with your own state and return early, for example if self._client is not None: return.
Added a link_format parameter to both PyPDFToDocument and PDFMinerToDocument components, matching the existing functionality in DOCXToDocument. Links are parsed from PDF annotations and appended at the bottom of the page content.
Added experimental context compaction for the Agent. CompactionHook runs before LLM calls and shortens the conversation when it reaches a configured fraction of the model's context window.
The first built-in strategy, SlidingWindowCompactor, preserves leading system messages, the latest user task, and as much complete recent conversation as the target allows. It removes complete historical turns first, and only when removing every historical turn is insufficient does it remove individual Agent steps from the current task. It replaces removed history with a short omission note, left where the removed messages used to sit: directly after the leading system messages when only historical turns were removed, and directly after the latest user message when the current task's own steps were removed. Only one note is ever present, because a later compaction folds an earlier one into its replacement.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.hooks.compaction import CompactionHook, SlidingWindowCompactor
hook = CompactionHook(
compactor=SlidingWindowCompactor(),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-nano"),
tools=[web_search],
hooks={"before_llm": [hook]},
)The hook uses provider-reported context usage when available and locally estimates the request otherwise, including tool schemas. Leave headroom above compact_at for the next reply and its tool results.
SlidingWindowCompactor treats an assistant message and its following tool results as one step, so a tool call is never separated from its results. Historical turns are likewise kept or removed in full, so an assistant reply is not retained without the user message it answers. It can also land above the requested target rather than under it, because leading system messages and the current task are never removed and min_keep_steps holds on to the newest Agent steps whatever their size, so a long system prompt or one large tool result can leave the conversation well over the target. Compaction is lossy: removed messages cannot be recovered or summarized by this strategy. Implement the Compactor protocol to provide a custom strategy.
CompactionHook and SlidingWindowCompactor emit an ExperimentalWarning and may change without a deprecation cycle.
Agent now returns an exit_reason output reporting why the run stopped, making it easier to route the Agent's output downstream (for example with a ConditionalRouter). It is one of: "text" (the model returned a reply with no tool calls), the name of the tool that satisfied a tool exit condition (in which case last_message is that tool's result), or "max_agent_steps" (the Agent hit max_agent_steps before meeting an exit condition). The reason is also available to hooks via state.get("exit_reason"), so an after_run hook can, for instance, append a fallback answer when the step budget is exhausted.
Add OpenAITokenCounter, which uses OpenAI's input token counting API to return model-specific counts for Haystack ChatMessage objects and optional tool schemas. Unlike local estimates, it supports OpenAI's exact accounting for request formatting, images, files, and tools.
Here is an example:
from haystack.dataclasses import ChatMessage
from haystack.token_counters import OpenAITokenCounter
counter = OpenAITokenCounter("gpt-5-mini")
count = counter.count([ChatMessage.from_user("Hello!")])Added haystack.token_counters: a TokenCounter protocol for estimating how many tokens a list of ChatMessage objects occupies, with two implementations.
Providers report token usage only after a call, and only for the call as a whole, so anything that needs a size beforehand - deciding whether a conversation still fits a model's context window, or how much of it to drop - has to estimate one.
from haystack.dataclasses import ChatMessage
from haystack.token_counters import ApproximateTokenCounter, TiktokenCounter
messages = [ChatMessage.from_user("Hello, how are you?")]
# No dependencies: estimates from text length.
ApproximateTokenCounter(chars_per_token=4.0).count(messages)
# Closer for OpenAI models; needs: pip install tiktoken
TiktokenCounter(encoding="o200k_base").count(messages)ApproximateTokenCounter needs nothing installed and estimates from text length at a configurable chars_per_token. TiktokenCounter counts with OpenAI's byte-pair encoder, which is closer for OpenAI models but requires tiktoken and drifts on other providers; it raises at construction when the dependency is missing, and loads its encoding on first use.
Neither can measure an image or a file, since a tokenizer only sees text and providers derive an image's cost from its dimensions. Both charge a flat tokens_per_image and tokens_per_file instead, counting images a tool returned inside its result as well as those a message carries directly. Raise those values if you send large images or long documents.
Tool schemas are sent alongside the messages and consume tokens too, so count takes an optional tools argument to have them included:
counter.count(messages, tools=[my_tool])Implement TokenCounter to count differently - for instance against a provider's own token-counting endpoint, which is the only way to have images counted exactly.
Added the experimental ToolResultPruningCompactor. It reduces Agent context usage by replacing older, large tool results with short placeholders while preserving tool-call/result structure. Results from a configurable number of recent tool-calling Agent steps remain intact, including parallel results from those steps.
from haystack.hooks.compaction import CompactionHook, ToolResultPruningCompactor
compaction_hook = CompactionHook(
compactor=ToolResultPruningCompactor(
min_keep_steps=2,
min_tokens=200,
),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)Add an Agent.clone() method that returns a new Agent with the same configuration, optionally replacing some init parameters: variant = agent.clone(system_prompt="Answer in German.").
Added AgentTool, a Tool that wraps a Haystack Agent, allowing it to be used as a tool by another Agent. It is a building block for multi-agent systems: an Agent specialized in one task becomes a tool that another Agent can delegate to. The calling Agent only sees the final reply, so all the steps the wrapped Agent takes stay out of its context. Sensible defaults make this work out of the box: the task is delegated as a single user message and comes back as text.
Example:
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import AgentTool, ComponentTool
from haystack_integrations.components.websearch.serperdev import SerperDevWebSearch
researcher = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-mini"),
system_prompt="You are a research specialist. Investigate the task and report your findings.",
tools=[
ComponentTool(
component=SerperDevWebSearch(
top_k=3,
),
name="web_search",
description="Search the web for current information on any topic",
),
],
)
research = AgentTool(
agent=researcher,
name="research",
description="Research a question on the web and report the findings",
)
coordinator = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4"),
tools=[research],
system_prompt="You coordinate specialists. Delegate research questions, then answer the user.",
)
result = coordinator.run([ChatMessage.from_user("What are the latest developments in the Haystack framework?")])
print(result["last_message"].text)Haystack components that use a document store now provide close and close_async methods for releasing resources. These methods are available on: AutoMergingRetriever, CacheChecker, DocumentWriter, FilterRetriever, and SentenceWindowRetriever. If the underlying Document Store does not implement the corresponding method, calling close or close_async has no effect.
Add a content-free haystack.agent.hook tracing span for every Agent hook invocation. Each span identifies the hook point, hook name, and hook type, allowing hook latency and failures to be attributed without tracing the potentially large Agent State. CompactionHook also adds its configured compaction strategy, estimated context size, whether compaction was triggered, its token target, and whether the compactor returned a replacement.
DOCXLinkFormat to a reusable LinkFormat Enum in haystack/components/converters/utils.py. DOCXLinkFormat is now an alias for backward compatibility.Agent now tracks an approximate current context-window size in its internal State under context_tokens, refreshed after every LLM call with that reply's prompt-plus-completion tokens (normalized across the prompt_tokens/completion_tokens and input_tokens/output_tokens key conventions). Unlike token_usage, which accumulates across the whole run, context_tokens is replaced each call. Hooks can read it via state.get("context_tokens") — for example, a before_llm hook that triggers context compaction once the value crosses a threshold. It is a best-effort snapshot: it is 0 when the generator does not report usage, and does not count messages appended after the latest call until the next call refreshes it.before_run hooks can now read state.data["tools"]. The key was previously written only once the first step had started, so a before_run hook hit a KeyError. It holds a snapshot of the tools available at that point, refreshed before every LLM call, so with a dynamic toolset such as SearchableToolset a before_run hook sees only the tools discovered so far.strict_datetime_comparison keyword argument to document_matches_filter, InMemoryDocumentStore, and MetadataRouter. When enabled, timezone-naive and timezone-aware datetimes never match each other. By default, mixed-awareness datetimes continue to be reconciled by copying the timezone from the aware value to the naive one, and this behavior is now consistent across equality, membership, and ordering operators.filters parameter to InMemoryDocumentStore.get_metadata_field_unique_values (sync and async), allowing the set of documents considered when computing unique metadata field values to be restricted.InMemoryDocumentStore.get_metadata_field_unique_values and its async counterpart now support pagination via from_ and size parameters, matching the behavior of other Document Stores (e.g. Chroma).MockChatGenerator's response_fn can now be tool-aware. If the callable accepts a second positional argument, it also receives the tools passed to run/run_async (a ToolsType or None), so a dynamic mock can build tool calls whose arguments follow the tool's parameter schema or route between the available tools. Existing single-argument response_fn callables are unaffected and keep receiving only the messages.State.to_dict now accepts a skip_keys parameter to exclude specific keys from the output.ConfirmationStrategy run to ConfirmationHook. Each haystack.agent.hook.human_in_the_loop.strategy span identifies the tool call it confirms and records the strategy type and the applied confirm, modify, or reject decision. When content tracing is enabled, spans also carry the arguments the strategy was run with and the ToolExecutionDecision it returned under haystack.agent.hook.human_in_the_loop.strategy.input and haystack.agent.hook.human_in_the_loop.strategy.output; chat messages, the confirmation strategy context, and the Agent State are not recorded by the hook.LLMEvaluator, LLMRanker, QueryExpander, LLMMetadataExtractor, LLMDocumentContentExtractor and LLMMessagesRouter now wrap their internal ChatGenerator calls in a haystack.chat_generator.run tracing span. These components do not return ChatMessage objects, so the LLM token usage carried in reply.meta["usage"] was previously lost to tracers. The new span exposes the generator's replies via the haystack.component.output tag, so token usage is now visible in traces (requires content tracing to be enabled). When the generator runs across threads, the span is nested under the component's span.+ operator (toolset_a + toolset_b) and passing a Toolset to add() (toolset_a.add(toolset_b)). Pass Toolsets as a list wherever tools are accepted instead: Agent(tools=[toolset_a, toolset_b]).Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). A malicious pipeline could either (a) set unsafe: true on an OutputAdapter or ConditionalRouter to disable the Jinja sandbox entirely, or (b) register the thread_safe_import import primitive as a Jinja custom_filters entry to import os and execute arbitrary commands — bypassing the deserialization allowlist and Jinja sandbox. The fix denies import primitives during callable deserialization, refuses to honor a component's unsafe flag while loading in safe mode, and hardens the Jinja sandbox (OutputAdapter, ConditionalRouter, PromptBuilder, ChatPromptBuilder) to block attribute access on module objects and calls into dangerous modules.Pipeline.loads / Pipeline.load / Pipeline.from_dict, which accept unsafe=True) and the execute primitives (Pipeline.run / run_async / run_async_generator / stream) were still resolvable from the allowlisted haystack namespace. Bound as a custom_filters entry on an OutputAdapter or ConditionalRouter (which bypass the Jinja sandbox), Pipeline.loads(..., unsafe=True) let a pipeline loaded in default safe mode load a nested pipeline whose own filters (allow_deserialization_module, deserialize_callable) bind under the nested unsafe context, disarming the process-wide allowlist with "*" and invoking os.system. All of these entry points are now marked as deserializer-internal, so they can never be produced by deserializing untrusted data. unsafe=True still bypasses the check by design; there are no public API changes.Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). Because the deserialization allowlist admits the whole haystack namespace, the deserializer's own allowlist-administration function (allow_deserialization_module) and its resolution helpers (deserialize_callable, deserialize_type, import_class_by_name) were themselves resolvable from serialized data. A malicious pipeline could register allow_deserialization_module as a Jinja custom_filters entry (on an OutputAdapter or ConditionalRouter), call it with "*" to disarm the allowlist process-wide, and then use the equally-resolvable deserialize_callable to resolve and invoke os.system. Loading alone was enough to trigger this: a Jinja filter called with constant arguments runs while the component is being constructed, so the pipeline never had to be run. The same attribute walk could also reach the deserializer's mutable control-plane state directly — for example a filter bound to _extra_allowed_modules.append — to widen the allowlist persistently and stage a later attack. Relatedly, the handle resolver walked attribute names freely, so a handle could descend into object internals such as <function>.__globals__ (a live module namespace, and via it __builtins__ and eval/exec) or <type>.__subclasses__ — classic sandbox-escape gadgets that stay inside an allowlisted module. The fix refuses to deserialize the deserialization control plane as a whole: the allowlist administration and resolution helpers (marked at definition time), everything defined in haystack.core.serialization_security, and any bound method of the mutable allowlist/context state. It also refuses to traverse into dunder and frame/code attributes while resolving a handle. This applies to both the callable- and class-resolution paths, and is bypassed only when the pipeline is loaded with unsafe=True.FileSystemToolResultStore.read() so it only reads references that resolve within the configured store root. This closes a boundary gap where callers could previously pass an arbitrary filesystem path to read() instead of a store-scoped reference returned by write().Fixed the serialization of PDFMinerToDocument. The component did not define to_dict, so the default serialization fell back to reading the init parameters from same-named attributes. Since the layout parameters are stored in self.layout_params, they were silently serialized with their default values, for example a component created with char_margin=0.5 was serialized with char_margin=2.0. Custom layout parameters are now preserved when a pipeline is serialized and loaded again.
Cancel and await sibling retrieval tasks when a concurrent call fails in MultiRetriever, MultiQueryTextRetriever, or MultiQueryEmbeddingRetriever.
Fixed an infinite recursion in CSVDocumentSplitter when nested row and column blocks were split together.
Comparing two Document objects with == now takes all metadata into account. Previously, two documents with different metadata could be considered equal if the metadata contained keys with the same names as document fields (such as id or content).
Document.from_dict(document.to_dict()) now correctly rebuilds any document. Previously, if the metadata contained keys with the same names as document fields (such as id or meta), this either raised an error or silently lost those metadata entries.
Fixed DocumentNDCGEvaluator producing NDCG scores outside the documented 0.0 to 1.0 range when the same document appeared more than once. A document retrieved multiple times used to be counted multiple times, pushing the score above 1.0; a ground truth document listed multiple times used to inflate the ideal gain, keeping a perfect retrieval below 1.0. Each distinct relevant document is now counted once, with the same relevance, in both the actual and ideal gain, so scores stay within range.
Fixed AnswerBuilder returning referenced documents in a scrambled order instead of ascending source-index order. The referenced document indices were collected in a set and iterated directly, so documents were emitted in the set's internal hash-table order (e.g. citations [3] [10] [50] yielded documents ordered 10, 3, 50). This order was deterministic but did not match the intuitive source order. The referenced documents are now returned sorted by their source index.
Fixed serialize_type and deserialize_type to correctly round-trip Callable types that declare an explicit parameter list, such as Callable[[int, str], bool]. Previously the parameter list was dropped during serialization (producing a malformed string like typing.Callable[, bool]) and could no longer be deserialized. This affected components that serialize type annotations, for example ConditionalRouter and OutputAdapter using a Callable output type.
Fix DocumentMAPEvaluator to include missed relevant documents in the average precision denominator and avoid crediting duplicate retrievals of the same document.
Fixed DocumentSplitter producing chunks that were not present in the source document when split_threshold was set together with split_overlap. Merging a below-threshold trailing segment into the previous split re-appended the overlapping units, duplicating text. The overlap is now added only once.
Fixed EmbeddingBasedDocumentSplitter.run_async embedding through the synchronous path while recursively splitting chunks longer than max_length. Only the first pass was async: the recursion called the sync splitting helper, so the embedder's blocking run ran on the event loop for every over-long chunk. The recursion now embeds through run_async as well.
Fixed JSONConverter raising a KeyError instead of logging its intended "Failed to extract text, skipping it" warning when a source is a ByteStream without a file_path in its meta (for example ByteStream.from_string(...), the exact usage shown in the component's own docstring examples). Affected error paths: invalid UTF-8 content, a jq_schema filter that fails to apply, and malformed JSON content.
Fixed LinkContentFetcher rotating the User-Agent on a cursor shared by every URL in the same run()/run_async() call. The URLs are fetched concurrently, so a retry triggered by one of them advanced the user agent for the others, and each completed fetch reset the cursor for the requests still in flight — most retries went out with the un-rotated user agent. Each fetch now walks the user_agents list on its own, so a URL rotates exactly as documented no matter how many other URLs are fetched alongside it.
Fixed MarkdownHeaderSplitter silently dropping a trailing header that has no body text. With keep_headers=True (the default), a header at the end of the document whose only content is whitespace was buffered to prepend to the next chunk, but with no following chunk it was never emitted, so the split documents no longer reconstructed the original text. Such trailing headers are now emitted as a final chunk.
Fixed MarkdownHeaderSplitter collapsing blank lines that follow a header with no body text. With keep_headers=True, such headers were re-joined with a single newline when prepended to the next chunk, so the split documents did not reconstruct the original text. Chunk content is now sliced from the original text and is byte-exact.
Fixed MarkdownHeaderSplitter including surrounding whitespace in the header and parent_headers metadata fields. The header text is now stripped; chunk content still keeps the header line's original whitespace.
Fixed schema-based serialization of lists, tuples and sets holding mixed types. Previously the schema was derived from the first element only, so deserializing such a value raised an AttributeError or silently returned mis-typed data (for example an Agent State field or a pipeline breakpoint input holding [Document(...), "text", 3]). Mixed-type arrays now record one schema per position using the JSON Schema prefixItems keyword and round-trip correctly. Homogeneous arrays keep the exact same output as before, so existing snapshots still load.
Fixed MSGToDocument raising a KeyError when converting a ByteStream source that has no file_path in its meta (for example a bare ByteStream(data=...), rather than a file path or a stream produced via ByteStream.from_file_path). Attachments extracted from such a source no longer include a parent_file_path key, since there is no source file path to record.
Pipeline connections now always convert values in the same way. When a component output is connected to an input that accepts multiple types, Haystack is sometimes able to automatically convert the value, and more than one conversion may be possible. For example, a ChatMessage with text "hello" connected to an input annotated str | list[str] can be delivered either as plain text ("hello") or as a list containing the text (["hello"]). Previously the conversion strategy was chosen non-deterministically, so the same pipeline could return a different value across runs. The conversion strategy is now selected using a fixed priority: first, wrapping a value in a list or unwrapping a single-element list; second, converting between ChatMessage and str; and last, combining both conversions.
Fixed normalize_metadata (used by all file converters) returning the same dictionary object for every source when meta is None or a single dictionary. Each source now receives an independent copy, so mutating one source's metadata downstream no longer leaks into the others.
OpenAIResponsesChatGenerator no longer mutates the parameters schema of the Tool objects passed to it. Previously every run wrote additionalProperties: False into the user's live Tool.parameters, silently altering the tool for any other generator that shared the same Tool instance and making serialization round trips unstable.
OpenAIResponsesChatGenerator no longer raises IndexError when it is warmed up with an empty tools list.
Fixed the parent of the haystack.agent.step.tool spans when an Agent step invokes several tools. The parent span is now resolved once before the tools run, so all tool calls of a step appear as siblings. Previously each span asked the tracer for the currently active span from inside its own concurrent invocation, which made the tool calls after the first one appear nested under a sibling tool call.
Fixed the haystack.pipeline.output_data tracing tag being empty. The tag was set at the start of Pipeline.run/run_async from the still-empty outputs, and since tracing backends coerce a tag value when it is set, the recorded output was always an empty dictionary. It is now set once the run completes so it reflects the final pipeline outputs. The tag is also gated behind content tracing (HAYSTACK_CONTENT_TRACING_ENABLED), consistent with the component-level input/output tags.
Fixed resuming a Pipeline from a pipeline_snapshot that was taken on a component's second or later visit, which failed with PipelineComponentsBlockedError: Cannot run pipeline - all components are blocked. A snapshot stored only the values of the pipeline's inputs and dropped the information about which component had sent each one, so on resume every restored input looked like it came from outside the pipeline, and such an input can only trigger a component on its first visit. Snapshots now record the sender of each input. Snapshots created by earlier versions of Haystack behave as before, so re-create them to resume anywhere in a looping pipeline.
Fixed a resumed Pipeline passing malformed inputs to the component the snapshot was taken on, whenever that component ran more than once after the resume, for example inside a loop. Every visit after the first reused the handling meant only for the paused visit and skipped the regular input consumption, so a variadic component could receive a bare value where it expected a list, raising errors such as TypeError: object of type 'int' has no len() from a BranchJoiner. This affected snapshots taken at any visit count, including the first.
Fixed QueryExpander returning duplicate queries when the chat generator repeats an expansion. Generated queries are now deduplicated while preserving first-seen order, so repeated expansions no longer trigger redundant retrievals or consume the requested expansion budget. Both run and run_async are affected.
RecursiveDocumentSplitter's word-mode fixed-size fallback no longer counts a run of whitespace (e.g. a double space, tab, or page break) as a word, so it no longer produces chunks smaller than split_length. It also no longer emits a whitespace-only chunk when the text ends in whitespace right after a chunk boundary; that trailing whitespace is now attached to the previous chunk instead.
This changes the exact chunk boundaries and chunk count produced by the word-unit fallback for any text containing such whitespace runs. Documents already split and indexed under the old behavior will produce different chunks if re-split after upgrading, so re-index any document store that relies on stable chunk boundaries from this fallback path.
Fixed an issue where PipelineBase.remove_component did not reset auto-variadic socket flags (is_lazy_variadic and wrap_input_in_list) on input sockets when components or connections were removed.
Fixed Pipeline.remove_component leaving dangling references to the removed component on the sockets of its neighboring components. Previously, removing a component reset only its own sockets, so a surviving neighbor kept the removed component's name in its input socket's senders (or output socket's receivers). This corrupted introspection and validation: Pipeline.inputs() hid a now-unconnected mandatory input, and feeding that input directly could raise a spurious "already connected" error. The removed component's name is now stripped from its neighbors' sockets as well.
SentenceWindowRetriever.run and SentenceWindowRetriever.run_async now validate an explicitly provided window_size=0 instead of treating it as unset and falling back to the constructor value.
Fixed the schema-aware serialization helper used for pipeline snapshots and Agent State (_serialize_value_with_schema) so it no longer silently passes unsupported objects through as if they were serialized. Values such as datetime, bytes, complex and arbitrary objects without a to_dict method were previously stored unchanged and mislabeled as strings, which broke JSON storage and round-tripping of snapshots. Unsupported values now raise a SerializationError, and the callers that build snapshots (pipeline breakpoints and State.to_dict) catch it to omit only the offending field while keeping the rest of the payload resumable.
Added support for serializing and deserializing frozenset values in _serialize_value_with_schema. A frozenset now round-trips back to a frozenset instead of being dropped.
Fixed serialize_type/deserialize_type for typing.Literal. Previously a Literal type hint was serialized with its values rendered as bare tokens (e.g. typing.Literal[yes, no]), which failed to deserialize, and values that looked like type names (e.g. Literal["int", "str"]) were silently turned into types on the round-trip. The values are now serialized with repr() and read back with ast.literal_eval, so a Literal type used by a component (such as OutputAdapter or ConditionalRouter) round-trips correctly through pipeline serialization.
Fixed an AttributeError: 'str' object has no attribute 'items' raised by create_tool_from_function, the @tool decorator, and ComponentTool when a tool parameter is named properties. Keys inside a JSON schema properties mapping are property names and are no longer misinterpreted as schema keywords when stripping the auto-generated title keywords.
Fixed create_tool_from_function, the @tool decorator, and ComponentTool corrupting a tool's JSON schema when the string title appears as a name rather than as a schema keyword. Stripping the auto-generated title keywords no longer deletes entries of $defs, definitions, patternProperties, dependentSchemas or dependentRequired (which would leave a $ref dangling or silently drop a validation rule), and no longer edits title keys inside default, const, enum or examples values, which are instance data and part of the tool's contract.
Fixed _ToolsetWrapper.__getitem__ (used when combining Toolsets with +) raising IndexError for negative indices, unlike a plain Toolset. Indexing a combined toolset now behaves consistently with a list of Tools, as documented.
Fixed serialization of types that contain ... (Ellipsis), such as variadic tuples (tuple[int, ...]) and Callable[..., X]. Previously serialize_type rendered the ... as the literal string "Ellipsis", which deserialize_type then rejected as a non-type builtin, so a component using such a type (for example OutputAdapter(output_type=tuple[int, ...])) could be serialized but not deserialized, breaking Pipeline.loads() / Pipeline.load(). These types now round-trip correctly, and pipelines serialized by older versions (which emitted "Ellipsis") can still be loaded.
Fixed ConfirmationHook applying a Human-in-the-Loop decision to the wrong tool call when a custom ConfirmationStrategy returns a decision with a missing or incorrect tool_call_id. Each decision is now bound to the tool call for which its strategy ran, and ID-bearing decisions are no longer matched to a different call by name. Haystack's existing requirement of exactly one decision per tool call is now explicitly enforced. Each matched decision is consumed after use, so a missing, unused, or reused decision raises a ValueError instead of being silently misapplied.
DocumentJoiner and AnswerJoiner now resolve top_k consistently and validate it. Previously, a runtime top_k=0 was treated as "unset" and silently fell back to the instance's top_k, instead of returning an empty list as requested. Both components now:
ValueError at initialization if top_k is not None and is less than or equal to 0.ValueError at runtime if top_k passed to run() is negative.run() is called with top_k=0, regardless of the instance's configured top_k.Fixes MetaFieldRanker silently treating a runtime top_k=0 as unset and falling back to the value configured at initialization. Runtime values that are not greater than zero now raise a ValueError as documented.
Fixed PythonCodeSplitter losing identifying context for oversized functions, methods, or classes. When a unit is too large and falls back to line-based secondary splitting, only the first resulting piece naturally retains the source def/class line; every piece now includes a qualified_name field in meta identifying the function, method, or class it came from.
Fixed RecursiveDocumentSplitter not setting the source_id meta field on the chunks it produces. It wrote only parent_id, while every other splitter in the library (DocumentSplitter, CSVDocumentSplitter, EmbeddingBasedDocumentSplitter, HierarchicalDocumentSplitter, MarkdownHeaderSplitter and PythonCodeSplitter) writes source_id. Components that follow that convention therefore rejected its output: SentenceWindowRetriever reads source_id by default and raises when it is absent, so it failed with "The retrieved documents must have 'source_id' in their metadata." on a pipeline that worked with any other splitter. Chunks now carry source_id as well as parent_id, which keeps its previous value for callers already reading it.
Tool functions defined in a module using from __future__ import annotations are now inspected correctly by Agent. Postponed annotations are stored as strings, so a parameter annotated with State was not recognized and the live State object was not injected into the tool call. The annotations are now resolved before they are inspected.
MarkdownHeaderSplitter and CSVDocumentSplitter now deep-copy the metadata of the document they split, matching DocumentSplitter. Previously they copied it shallowly, so nested values such as a list under meta["tags"] were shared between every chunk and with the input document, and editing one chunk's metadata changed all the others. HierarchicalDocumentSplitter had the same problem on its root node, which kept references into the input document's metadata.
Keep an Agent's execution counter in sync with step_count restored by a before_run hook, so restarted Agents continue from the saved step instead of resetting the count.
@Aarkin7, @anakin87, @anxkhn, @aquib8112, @Aryan-Pardeshi, @atikulmunna, @bharadwaj-pendyala, @bilgeyucel, @bogdankostic, @camgrimsec, @chuenchen309, @davidpavlovschi, @davidsbatista, @DhanushPillay, @DivyaNarahari97, @erikos, @GovindhKishore, @hxaxd, @immuhammadfurqan, @iridescentWen, @jaideeppyne, @julian-risch, @kacperlukawski, @KXHXK, @LHMQ878, @LK-maker-007, @lntutor, @manjunathbhaskar, @mittalpk, @MVS-source, @onatozmenn, @otiscuilei, @pcbeingused333, @rautaditya2606, @sjrl, @sohumt123, @Solaris-star, @TimurRakhmatullin86, @vidigoat, @vinkiYu, @winklemad, @yaodong-shen
Fixed a remote code execution vulnerability that could be triggered by loading an untrusted pipeline in default safe mode ( Pipeline.load / Pipeline.l…
exit_reason is now a reserved state key on Agent. If you defined a custom state_schema key named exit_reason, rename it: the Agent now raises a ValueError at initialization when a reserved key is redefined.
Agent.state_schema now contains the user-provided state schema, exactly as passed to __init__. Previously it contained the resolved schema, which also included messages and the keys the Agent manages internally (step_count, token_usage, exit_reason, ...). If you were reading agent.state_schema to inspect the effective runtime schema, use the new public attribute agent.resolved_state_schema instead.
DocumentMAPEvaluator scores can change because average precision now uses all unique, valid ground-truth comparison values as its denominator and credits each value at most once. Re-baseline evaluations that relied on the previous scores.
PipelineSnapshot.pipeline_state.inputs changed shape. It used to store one flattened value per socket, {component: {socket: value}}. It now stores the pipeline's internal inputs, keeping the component that sent each one in the order it arrived: {component: {socket: [{"sender": ..., "value": ...}]}}. PipelineState gained an inputs_format field recording which of the two shapes a snapshot uses. The same applies to BreakpointException.inputs, which returns that field.
You are affected if you read pipeline_state.inputs (or BreakpointException.inputs) directly, for example to display or post-process a snapshot. Resuming a pipeline with Pipeline.run(pipeline_snapshot=...) is not affected, and neither is code that only passes snapshots around or persists them.
To adapt, read the input from the list and take its value:
inputs = snapshot.pipeline_state.inputs["serialized_data"]
# before
value = inputs["my_component"]["my_socket"]
# now
value = inputs["my_component"]["my_socket"][0]["value"]A socket that received inputs from several senders has one list item per input, each recording the sender that produced it.
Snapshots written by earlier versions of Haystack have inputs_format set to None and keep the flattened shape, so branch on that field if you need to handle both.
Loading a serialized OutputAdapter or ConditionalRouter whose unsafe init parameter is set to true now raises DeserializationError unless the pipeline is loaded in unsafe mode. Pipelines that legitimately rely on an unsafe OutputAdapter/ConditionalRouter embedded in serialized data must now load with Pipeline.load(..., unsafe=True) (or Pipeline.loads / Pipeline.from_dict with unsafe=True).
Passing window_size=0 to SentenceWindowRetriever.run or SentenceWindowRetriever.run_async now raises a ValueError instead of silently using the window_size set in the constructor. You are affected if you pass window_size=0 at runtime, either directly or from an upstream component in a pipeline. If you were relying on 0 to mean "use the value from the constructor", omit the argument (or pass None) instead:
retriever = SentenceWindowRetriever(document_store=document_store, window_size=3)
# Before: silently used window_size=3
retriever.run(retrieved_documents=docs, window_size=0)
# After: omit the argument to use the constructor value
retriever.run(retrieved_documents=docs)InMemoryDocumentStore.get_metadata_field_unique_values (and its async counterpart)'s search_term parameter now matches against the metadata field's own value (case-insensitive substring) instead of the document's content. Callers relying on the previous content-matching behavior will need to filter documents by content themselves before calling this method.
The Agent now warms up its hooks before every run, not only the first one, as it already does for Tools and Toolsets. If your hook has a warm_up() that does expensive setup (opening a client, loading a model), make it return early once done, for example if self._client is not None: return.
Haystack can call warm_up() on Tools and Toolsets more than once, for example before every run. Previously Toolset absorbed repeated calls with an internal _is_warmed_up flag; that flag is gone and every call now reaches your warm_up(). If your custom Tool or Toolset does expensive work there (connecting to a server, loading a model), or relied on the _is_warmed_up attribute, guard with your own state and return early, for example if self._client is not None: return.
Added a link_format parameter to both PyPDFToDocument and PDFMinerToDocument components, matching the existing functionality in DOCXToDocument. Links are parsed from PDF annotations and appended at the bottom of the page content.
Added experimental context compaction for the Agent. CompactionHook runs before LLM calls and shortens the conversation when it reaches a configured fraction of the model's context window.
The first built-in strategy, SlidingWindowCompactor, preserves leading system messages, the latest user task, and as much complete recent conversation as the target allows. It removes complete historical turns first, and only when removing every historical turn is insufficient does it remove individual Agent steps from the current task. It replaces removed history with a short omission note, left where the removed messages used to sit: directly after the leading system messages when only historical turns were removed, and directly after the latest user message when the current task's own steps were removed. Only one note is ever present, because a later compaction folds an earlier one into its replacement.
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.hooks.compaction import CompactionHook, SlidingWindowCompactor
hook = CompactionHook(
compactor=SlidingWindowCompactor(),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)
agent = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-nano"),
tools=[web_search],
hooks={"before_llm": [hook]},
)The hook uses provider-reported context usage when available and locally estimates the request otherwise, including tool schemas. Leave headroom above compact_at for the next reply and its tool results.
SlidingWindowCompactor treats an assistant message and its following tool results as one step, so a tool call is never separated from its results. Historical turns are likewise kept or removed in full, so an assistant reply is not retained without the user message it answers. It can also land above the requested target rather than under it, because leading system messages and the current task are never removed and min_keep_steps holds on to the newest Agent steps whatever their size, so a long system prompt or one large tool result can leave the conversation well over the target. Compaction is lossy: removed messages cannot be recovered or summarized by this strategy. Implement the Compactor protocol to provide a custom strategy.
CompactionHook and SlidingWindowCompactor emit an ExperimentalWarning and may change without a deprecation cycle.
Agent now returns an exit_reason output reporting why the run stopped, making it easier to route the Agent's output downstream (for example with a ConditionalRouter). It is one of: "text" (the model returned a reply with no tool calls), the name of the tool that satisfied a tool exit condition (in which case last_message is that tool's result), or "max_agent_steps" (the Agent hit max_agent_steps before meeting an exit condition). The reason is also available to hooks via state.get("exit_reason"), so an after_run hook can, for instance, append a fallback answer when the step budget is exhausted.
Add OpenAITokenCounter, which uses OpenAI's input token counting API to return model-specific counts for Haystack ChatMessage objects and optional tool schemas. Unlike local estimates, it supports OpenAI's exact accounting for request formatting, images, files, and tools.
Here is an example:
from haystack.dataclasses import ChatMessage
from haystack.token_counters import OpenAITokenCounter
counter = OpenAITokenCounter("gpt-5-mini")
count = counter.count([ChatMessage.from_user("Hello!")])Added haystack.token_counters: a TokenCounter protocol for estimating how many tokens a list of ChatMessage objects occupies, with two implementations.
Providers report token usage only after a call, and only for the call as a whole, so anything that needs a size beforehand - deciding whether a conversation still fits a model's context window, or how much of it to drop - has to estimate one.
from haystack.dataclasses import ChatMessage
from haystack.token_counters import ApproximateTokenCounter, TiktokenCounter
messages = [ChatMessage.from_user("Hello, how are you?")]
# No dependencies: estimates from text length.
ApproximateTokenCounter(chars_per_token=4.0).count(messages)
# Closer for OpenAI models; needs: pip install tiktoken
TiktokenCounter(encoding="o200k_base").count(messages)ApproximateTokenCounter needs nothing installed and estimates from text length at a configurable chars_per_token. TiktokenCounter counts with OpenAI's byte-pair encoder, which is closer for OpenAI models but requires tiktoken and drifts on other providers; it raises at construction when the dependency is missing, and loads its encoding on first use.
Neither can measure an image or a file, since a tokenizer only sees text and providers derive an image's cost from its dimensions. Both charge a flat tokens_per_image and tokens_per_file instead, counting images a tool returned inside its result as well as those a message carries directly. Raise those values if you send large images or long documents.
Tool schemas are sent alongside the messages and consume tokens too, so count takes an optional tools argument to have them included:
counter.count(messages, tools=[my_tool])Implement TokenCounter to count differently - for instance against a provider's own token-counting endpoint, which is the only way to have images counted exactly.
Added the experimental ToolResultPruningCompactor. It reduces Agent context usage by replacing older, large tool results with short placeholders while preserving tool-call/result structure. Results from a configurable number of recent tool-calling Agent steps remain intact, including parallel results from those steps.
from haystack.hooks.compaction import CompactionHook, ToolResultPruningCompactor
compaction_hook = CompactionHook(
compactor=ToolResultPruningCompactor(
min_keep_steps=2,
min_tokens=200,
),
context_window=400_000,
compact_at=0.7,
compact_to=0.4,
)Add an Agent.clone() method that returns a new Agent with the same configuration, optionally replacing some init parameters: variant = agent.clone(system_prompt="Answer in German.").
Added AgentTool, a Tool that wraps a Haystack Agent, allowing it to be used as a tool by another Agent. It is a building block for multi-agent systems: an Agent specialized in one task becomes a tool that another Agent can delegate to. The calling Agent only sees the final reply, so all the steps the wrapped Agent takes stay out of its context. Sensible defaults make this work out of the box: the task is delegated as a single user message and comes back as text.
Example:
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIResponsesChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import AgentTool, ComponentTool
from haystack_integrations.components.websearch.serperdev import SerperDevWebSearch
researcher = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4-mini"),
system_prompt="You are a research specialist. Investigate the task and report your findings.",
tools=[
ComponentTool(
component=SerperDevWebSearch(
top_k=3,
),
name="web_search",
description="Search the web for current information on any topic",
),
],
)
research = AgentTool(
agent=researcher,
name="research",
description="Research a question on the web and report the findings",
)
coordinator = Agent(
chat_generator=OpenAIResponsesChatGenerator(model="gpt-5.4"),
tools=[research],
system_prompt="You coordinate specialists. Delegate research questions, then answer the user.",
)
result = coordinator.run([ChatMessage.from_user("What are the latest developments in the Haystack framework?")])
print(result["last_message"].text)Haystack components that use a document store now provide close and close_async methods for releasing resources. These methods are available on: AutoMergingRetriever, CacheChecker, DocumentWriter, FilterRetriever, and SentenceWindowRetriever. If the underlying Document Store does not implement the corresponding method, calling close or close_async has no effect.
Add a content-free haystack.agent.hook tracing span for every Agent hook invocation. Each span identifies the hook point, hook name, and hook type, allowing hook latency and failures to be attributed without tracing the potentially large Agent State. CompactionHook also adds its configured compaction strategy, estimated context size, whether compaction was triggered, its token target, and whether the compactor returned a replacement.
DOCXLinkFormat to a reusable LinkFormat Enum in haystack/components/converters/utils.py. DOCXLinkFormat is now an alias for backward compatibility.Agent now tracks an approximate current context-window size in its internal State under context_tokens, refreshed after every LLM call with that reply's prompt-plus-completion tokens (normalized across the prompt_tokens/completion_tokens and input_tokens/output_tokens key conventions). Unlike token_usage, which accumulates across the whole run, context_tokens is replaced each call. Hooks can read it via state.get("context_tokens") — for example, a before_llm hook that triggers context compaction once the value crosses a threshold. It is a best-effort snapshot: it is 0 when the generator does not report usage, and does not count messages appended after the latest call until the next call refreshes it.before_run hooks can now read state.data["tools"]. The key was previously written only once the first step had started, so a before_run hook hit a KeyError. It holds a snapshot of the tools available at that point, refreshed before every LLM call, so with a dynamic toolset such as SearchableToolset a before_run hook sees only the tools discovered so far.strict_datetime_comparison keyword argument to document_matches_filter, InMemoryDocumentStore, and MetadataRouter. When enabled, timezone-naive and timezone-aware datetimes never match each other. By default, mixed-awareness datetimes continue to be reconciled by copying the timezone from the aware value to the naive one, and this behavior is now consistent across equality, membership, and ordering operators.filters parameter to InMemoryDocumentStore.get_metadata_field_unique_values (sync and async), allowing the set of documents considered when computing unique metadata field values to be restricted.InMemoryDocumentStore.get_metadata_field_unique_values and its async counterpart now support pagination via from_ and size parameters, matching the behavior of other Document Stores (e.g. Chroma).MockChatGenerator's response_fn can now be tool-aware. If the callable accepts a second positional argument, it also receives the tools passed to run/run_async (a ToolsType or None), so a dynamic mock can build tool calls whose arguments follow the tool's parameter schema or route between the available tools. Existing single-argument response_fn callables are unaffected and keep receiving only the messages.State.to_dict now accepts a skip_keys parameter to exclude specific keys from the output.ConfirmationStrategy run to ConfirmationHook. Each haystack.agent.hook.human_in_the_loop.strategy span identifies the tool call it confirms and records the strategy type and the applied confirm, modify, or reject decision. When content tracing is enabled, spans also carry the arguments the strategy was run with and the ToolExecutionDecision it returned under haystack.agent.hook.human_in_the_loop.strategy.input and haystack.agent.hook.human_in_the_loop.strategy.output; chat messages, the confirmation strategy context, and the Agent State are not recorded by the hook.LLMEvaluator, LLMRanker, QueryExpander, LLMMetadataExtractor, LLMDocumentContentExtractor and LLMMessagesRouter now wrap their internal ChatGenerator calls in a haystack.chat_generator.run tracing span. These components do not return ChatMessage objects, so the LLM token usage carried in reply.meta["usage"] was previously lost to tracers. The new span exposes the generator's replies via the haystack.component.output tag, so token usage is now visible in traces (requires content tracing to be enabled). When the generator runs across threads, the span is nested under the component's span.+ operator (toolset_a + toolset_b) and passing a Toolset to add() (toolset_a.add(toolset_b)). Pass Toolsets as a list wherever tools are accepted instead: Agent(tools=[toolset_a, toolset_b]).Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). A malicious pipeline could either (a) set unsafe: true on an OutputAdapter or ConditionalRouter to disable the Jinja sandbox entirely, or (b) register the thread_safe_import import primitive as a Jinja custom_filters entry to import os and execute arbitrary commands — bypassing the deserialization allowlist and Jinja sandbox. The fix denies import primitives during callable deserialization, refuses to honor a component's unsafe flag while loading in safe mode, and hardens the Jinja sandbox (OutputAdapter, ConditionalRouter, PromptBuilder, ChatPromptBuilder) to block attribute access on module objects and calls into dangerous modules.Pipeline.loads / Pipeline.load / Pipeline.from_dict, which accept unsafe=True) and the execute primitives (Pipeline.run / run_async / run_async_generator / stream) were still resolvable from the allowlisted haystack namespace. Bound as a custom_filters entry on an OutputAdapter or ConditionalRouter (which bypass the Jinja sandbox), Pipeline.loads(..., unsafe=True) let a pipeline loaded in default safe mode load a nested pipeline whose own filters (allow_deserialization_module, deserialize_callable) bind under the nested unsafe context, disarming the process-wide allowlist with "*" and invoking os.system. All of these entry points are now marked as deserializer-internal, so they can never be produced by deserializing untrusted data. unsafe=True still bypasses the check by design; there are no public API changes.Pipeline.load / Pipeline.loads / Pipeline.from_dict, without unsafe=True). Because the deserialization allowlist admits the whole haystack namespace, the deserializer's own allowlist-administration function (allow_deserialization_module) and its resolution helpers (deserialize_callable, deserialize_type, import_class_by_name) were themselves resolvable from serialized data. A malicious pipeline could register allow_deserialization_module as a Jinja custom_filters entry (on an OutputAdapter or ConditionalRouter), call it with "*" to disarm the allowlist process-wide, and then use the equally-resolvable deserialize_callable to resolve and invoke os.system. Loading alone was enough to trigger this: a Jinja filter called with constant arguments runs while the component is being constructed, so the pipeline never had to be run. The same attribute walk could also reach the deserializer's mutable control-plane state directly — for example a filter bound to _extra_allowed_modules.append — to widen the allowlist persistently and stage a later attack. Relatedly, the handle resolver walked attribute names freely, so a handle could descend into object internals such as <function>.__globals__ (a live module namespace, and via it __builtins__ and eval/exec) or <type>.__subclasses__ — classic sandbox-escape gadgets that stay inside an allowlisted module. The fix refuses to deserialize the deserialization control plane as a whole: the allowlist administration and resolution helpers (marked at definition time), everything defined in haystack.core.serialization_security, and any bound method of the mutable allowlist/context state. It also refuses to traverse into dunder and frame/code attributes while resolving a handle. This applies to both the callable- and class-resolution paths, and is bypassed only when the pipeline is loaded with unsafe=True.FileSystemToolResultStore.read() so it only reads references that resolve within the configured store root. This closes a boundary gap where callers could previously pass an arbitrary filesystem path to read() instead of a store-scoped reference returned by write().Fixed the serialization of PDFMinerToDocument. The component did not define to_dict, so the default serialization fell back to reading the init parameters from same-named attributes. Since the layout parameters are stored in self.layout_params, they were silently serialized with their default values, for example a component created with char_margin=0.5 was serialized with char_margin=2.0. Custom layout parameters are now preserved when a pipeline is serialized and loaded again.
Cancel and await sibling retrieval tasks when a concurrent call fails in MultiRetriever, MultiQueryTextRetriever, or MultiQueryEmbeddingRetriever.
Fixed an infinite recursion in CSVDocumentSplitter when nested row and column blocks were split together.
Comparing two Document objects with == now takes all metadata into account. Previously, two documents with different metadata could be considered equal if the metadata contained keys with the same names as document fields (such as id or content).
Document.from_dict(document.to_dict()) now correctly rebuilds any document. Previously, if the metadata contained keys with the same names as document fields (such as id or meta), this either raised an error or silently lost those metadata entries.
Fixed DocumentNDCGEvaluator producing NDCG scores outside the documented 0.0 to 1.0 range when the same document appeared more than once. A document retrieved multiple times used to be counted multiple times, pushing the score above 1.0; a ground truth document listed multiple times used to inflate the ideal gain, keeping a perfect retrieval below 1.0. Each distinct relevant document is now counted once, with the same relevance, in both the actual and ideal gain, so scores stay within range.
Fixed AnswerBuilder returning referenced documents in a scrambled order instead of ascending source-index order. The referenced document indices were collected in a set and iterated directly, so documents were emitted in the set's internal hash-table order (e.g. citations [3] [10] [50] yielded documents ordered 10, 3, 50). This order was deterministic but did not match the intuitive source order. The referenced documents are now returned sorted by their source index.
Fixed serialize_type and deserialize_type to correctly round-trip Callable types that declare an explicit parameter list, such as Callable[[int, str], bool]. Previously the parameter list was dropped during serialization (producing a malformed string like typing.Callable[, bool]) and could no longer be deserialized. This affected components that serialize type annotations, for example ConditionalRouter and OutputAdapter using a Callable output type.
Fix DocumentMAPEvaluator to include missed relevant documents in the average precision denominator and avoid crediting duplicate retrievals of the same document.
Fixed DocumentSplitter producing chunks that were not present in the source document when split_threshold was set together with split_overlap. Merging a below-threshold trailing segment into the previous split re-appended the overlapping units, duplicating text. The overlap is now added only once.
Fixed EmbeddingBasedDocumentSplitter.run_async embedding through the synchronous path while recursively splitting chunks longer than max_length. Only the first pass was async: the recursion called the sync splitting helper, so the embedder's blocking run ran on the event loop for every over-long chunk. The recursion now embeds through run_async as well.
Fixed JSONConverter raising a KeyError instead of logging its intended "Failed to extract text, skipping it" warning when a source is a ByteStream without a file_path in its meta (for example ByteStream.from_string(...), the exact usage shown in the component's own docstring examples). Affected error paths: invalid UTF-8 content, a jq_schema filter that fails to apply, and malformed JSON content.
Fixed LinkContentFetcher rotating the User-Agent on a cursor shared by every URL in the same run()/run_async() call. The URLs are fetched concurrently, so a retry triggered by one of them advanced the user agent for the others, and each completed fetch reset the cursor for the requests still in flight — most retries went out with the un-rotated user agent. Each fetch now walks the user_agents list on its own, so a URL rotates exactly as documented no matter how many other URLs are fetched alongside it.
Fixed MarkdownHeaderSplitter silently dropping a trailing header that has no body text. With keep_headers=True (the default), a header at the end of the document whose only content is whitespace was buffered to prepend to the next chunk, but with no following chunk it was never emitted, so the split documents no longer reconstructed the original text. Such trailing headers are now emitted as a final chunk.
Fixed MarkdownHeaderSplitter collapsing blank lines that follow a header with no body text. With keep_headers=True, such headers were re-joined with a single newline when prepended to the next chunk, so the split documents did not reconstruct the original text. Chunk content is now sliced from the original text and is byte-exact.
Fixed MarkdownHeaderSplitter including surrounding whitespace in the header and parent_headers metadata fields. The header text is now stripped; chunk content still keeps the header line's original whitespace.
Fixed schema-based serialization of lists, tuples and sets holding mixed types. Previously the schema was derived from the first element only, so deserializing such a value raised an AttributeError or silently returned mis-typed data (for example an Agent State field or a pipeline breakpoint input holding [Document(...), "text", 3]). Mixed-type arrays now record one schema per position using the JSON Schema prefixItems keyword and round-trip correctly. Homogeneous arrays keep the exact same output as before, so existing snapshots still load.
Fixed MSGToDocument raising a KeyError when converting a ByteStream source that has no file_path in its meta (for example a bare ByteStream(data=...), rather than a file path or a stream produced via ByteStream.from_file_path). Attachments extracted from such a source no longer include a parent_file_path key, since there is no source file path to record.
Pipeline connections now always convert values in the same way. When a component output is connected to an input that accepts multiple types, Haystack is sometimes able to automatically convert the value, and more than one conversion may be possible. For example, a ChatMessage with text "hello" connected to an input annotated str | list[str] can be delivered either as plain text ("hello") or as a list containing the text (["hello"]). Previously the conversion strategy was chosen non-deterministically, so the same pipeline could return a different value across runs. The conversion strategy is now selected using a fixed priority: first, wrapping a value in a list or unwrapping a single-element list; second, converting between ChatMessage and str; and last, combining both conversions.
Fixed normalize_metadata (used by all file converters) returning the same dictionary object for every source when meta is None or a single dictionary. Each source now receives an independent copy, so mutating one source's metadata downstream no longer leaks into the others.
OpenAIResponsesChatGenerator no longer mutates the parameters schema of the Tool objects passed to it. Previously every run wrote additionalProperties: False into the user's live Tool.parameters, silently altering the tool for any other generator that shared the same Tool instance and making serialization round trips unstable.
OpenAIResponsesChatGenerator no longer raises IndexError when it is warmed up with an empty tools list.
Fixed the parent of the haystack.agent.step.tool spans when an Agent step invokes several tools. The parent span is now resolved once before the tools run, so all tool calls of a step appear as siblings. Previously each span asked the tracer for the currently active span from inside its own concurrent invocation, which made the tool calls after the first one appear nested under a sibling tool call.
Fixed the haystack.pipeline.output_data tracing tag being empty. The tag was set at the start of Pipeline.run/run_async from the still-empty outputs, and since tracing backends coerce a tag value when it is set, the recorded output was always an empty dictionary. It is now set once the run completes so it reflects the final pipeline outputs. The tag is also gated behind content tracing (HAYSTACK_CONTENT_TRACING_ENABLED), consistent with the component-level input/output tags.
Fixed resuming a Pipeline from a pipeline_snapshot that was taken on a component's second or later visit, which failed with PipelineComponentsBlockedError: Cannot run pipeline - all components are blocked. A snapshot stored only the values of the pipeline's inputs and dropped the information about which component had sent each one, so on resume every restored input looked like it came from outside the pipeline, and such an input can only trigger a component on its first visit. Snapshots now record the sender of each input. Snapshots created by earlier versions of Haystack behave as before, so re-create them to resume anywhere in a looping pipeline.
Fixed a resumed Pipeline passing malformed inputs to the component the snapshot was taken on, whenever that component ran more than once after the resume, for example inside a loop. Every visit after the first reused the handling meant only for the paused visit and skipped the regular input consumption, so a variadic component could receive a bare value where it expected a list, raising errors such as TypeError: object of type 'int' has no len() from a BranchJoiner. This affected snapshots taken at any visit count, including the first.
Fixed QueryExpander returning duplicate queries when the chat generator repeats an expansion. Generated queries are now deduplicated while preserving first-seen order, so repeated expansions no longer trigger redundant retrievals or consume the requested expansion budget. Both run and run_async are affected.
RecursiveDocumentSplitter's word-mode fixed-size fallback no longer counts a run of whitespace (e.g. a double space, tab, or page break) as a word, so it no longer produces chunks smaller than split_length. It also no longer emits a whitespace-only chunk when the text ends in whitespace right after a chunk boundary; that trailing whitespace is now attached to the previous chunk instead.
This changes the exact chunk boundaries and chunk count produced by the word-unit fallback for any text containing such whitespace runs. Documents already split and indexed under the old behavior will produce different chunks if re-split after upgrading, so re-index any document store that relies on stable chunk boundaries from this fallback path.
Fixed an issue where PipelineBase.remove_component did not reset auto-variadic socket flags (is_lazy_variadic and wrap_input_in_list) on input sockets when components or connections were removed.
Fixed Pipeline.remove_component leaving dangling references to the removed component on the sockets of its neighboring components. Previously, removing a component reset only its own sockets, so a surviving neighbor kept the removed component's name in its input socket's senders (or output socket's receivers). This corrupted introspection and validation: Pipeline.inputs() hid a now-unconnected mandatory input, and feeding that input directly could raise a spurious "already connected" error. The removed component's name is now stripped from its neighbors' sockets as well.
SentenceWindowRetriever.run and SentenceWindowRetriever.run_async now validate an explicitly provided window_size=0 instead of treating it as unset and falling back to the constructor value.
Fixed the schema-aware serialization helper used for pipeline snapshots and Agent State (_serialize_value_with_schema) so it no longer silently passes unsupported objects through as if they were serialized. Values such as datetime, bytes, complex and arbitrary objects without a to_dict method were previously stored unchanged and mislabeled as strings, which broke JSON storage and round-tripping of snapshots. Unsupported values now raise a SerializationError, and the callers that build snapshots (pipeline breakpoints and State.to_dict) catch it to omit only the offending field while keeping the rest of the payload resumable.
Added support for serializing and deserializing frozenset values in _serialize_value_with_schema. A frozenset now round-trips back to a frozenset instead of being dropped.
Fixed serialize_type/deserialize_type for typing.Literal. Previously a Literal type hint was serialized with its values rendered as bare tokens (e.g. typing.Literal[yes, no]), which failed to deserialize, and values that looked like type names (e.g. Literal["int", "str"]) were silently turned into types on the round-trip. The values are now serialized with repr() and read back with ast.literal_eval, so a Literal type used by a component (such as OutputAdapter or ConditionalRouter) round-trips correctly through pipeline serialization.
Fixed an AttributeError: 'str' object has no attribute 'items' raised by create_tool_from_function, the @tool decorator, and ComponentTool when a tool parameter is named properties. Keys inside a JSON schema properties mapping are property names and are no longer misinterpreted as schema keywords when stripping the auto-generated title keywords.
Fixed create_tool_from_function, the @tool decorator, and ComponentTool corrupting a tool's JSON schema when the string title appears as a name rather than as a schema keyword. Stripping the auto-generated title keywords no longer deletes entries of $defs, definitions, patternProperties, dependentSchemas or dependentRequired (which would leave a $ref dangling or silently drop a validation rule), and no longer edits title keys inside default, const, enum or examples values, which are instance data and part of the tool's contract.
Fixed _ToolsetWrapper.__getitem__ (used when combining Toolsets with +) raising IndexError for negative indices, unlike a plain Toolset. Indexing a combined toolset now behaves consistently with a list of Tools, as documented.
Fixed serialization of types that contain ... (Ellipsis), such as variadic tuples (tuple[int, ...]) and Callable[..., X]. Previously serialize_type rendered the ... as the literal string "Ellipsis", which deserialize_type then rejected as a non-type builtin, so a component using such a type (for example OutputAdapter(output_type=tuple[int, ...])) could be serialized but not deserialized, breaking Pipeline.loads() / Pipeline.load(). These types now round-trip correctly, and pipelines serialized by older versions (which emitted "Ellipsis") can still be loaded.
Fixed ConfirmationHook applying a Human-in-the-Loop decision to the wrong tool call when a custom ConfirmationStrategy returns a decision with a missing or incorrect tool_call_id. Each decision is now bound to the tool call for which its strategy ran, and ID-bearing decisions are no longer matched to a different call by name. Haystack's existing requirement of exactly one decision per tool call is now explicitly enforced. Each matched decision is consumed after use, so a missing, unused, or reused decision raises a ValueError instead of being silently misapplied.
DocumentJoiner and AnswerJoiner now resolve top_k consistently and validate it. Previously, a runtime top_k=0 was treated as "unset" and silently fell back to the instance's top_k, instead of returning an empty list as requested. Both components now:
ValueError at initialization if top_k is not None and is less than or equal to 0.ValueError at runtime if top_k passed to run() is negative.run() is called with top_k=0, regardless of the instance's configured top_k.Fixes MetaFieldRanker silently treating a runtime top_k=0 as unset and falling back to the value configured at initialization. Runtime values that are not greater than zero now raise a ValueError as documented.
Fixed PythonCodeSplitter losing identifying context for oversized functions, methods, or classes. When a unit is too large and falls back to line-based secondary splitting, only the first resulting piece naturally retains the source def/class line; every piece now includes a qualified_name field in meta identifying the function, method, or class it came from.
Fixed RecursiveDocumentSplitter not setting the source_id meta field on the chunks it produces. It wrote only parent_id, while every other splitter in the library (DocumentSplitter, CSVDocumentSplitter, EmbeddingBasedDocumentSplitter, HierarchicalDocumentSplitter, MarkdownHeaderSplitter and PythonCodeSplitter) writes source_id. Components that follow that convention therefore rejected its output: SentenceWindowRetriever reads source_id by default and raises when it is absent, so it failed with "The retrieved documents must have 'source_id' in their metadata." on a pipeline that worked with any other splitter. Chunks now carry source_id as well as parent_id, which keeps its previous value for callers already reading it.
Tool functions defined in a module using from __future__ import annotations are now inspected correctly by Agent. Postponed annotations are stored as strings, so a parameter annotated with State was not recognized and the live State object was not injected into the tool call. The annotations are now resolved before they are inspected.
MarkdownHeaderSplitter and CSVDocumentSplitter now deep-copy the metadata of the document they split, matching DocumentSplitter. Previously they copied it shallowly, so nested values such as a list under meta["tags"] were shared between every chunk and with the input document, and editing one chunk's metadata changed all the others. HierarchicalDocumentSplitter had the same problem on its root node, which kept references into the input document's metadata.
Keep an Agent's execution counter in sync with step_count restored by a before_run hook, so restarted Agents continue from the saved step instead of resetting the count.
@Aarkin7, @anakin87, @anxkhn, @aquib8112, @Aryan-Pardeshi, @atikulmunna, @bharadwaj-pendyala, @bilgeyucel, @bogdankostic, @camgrimsec, @chuenchen309, @davidpavlovschi, @davidsbatista, @DhanushPillay, @DivyaNarahari97, @erikos, @GovindhKishore, @hxaxd, @immuhammadfurqan, @iridescentWen, @jaideeppyne, @julian-risch, @kacperlukawski, @KXHXK, @LHMQ878, @LK-maker-007, @lntutor, @manjunathbhaskar, @mittalpk, @MVS-source, @onatozmenn, @otiscuilei, @pcbeingused333, @rautaditya2606, @sjrl, @sohumt123, @Solaris-star, @TimurRakhmatullin86, @vidigoat, @vinkiYu, @winklemad, @yaodong-shen
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →