NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2020 most downloaded on PyPI
A library that prepares raw documents for downstream ML tasks.
Last release 4 days ago
14 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
3 versions withdrawn
withdrawn after publishing
4 years old
233 releases · first in 2022
One column per quarter.
Refactoring the VoyageAI integration to use voyageai package directly, allowing extra features.
build_layout_elements_from_cor_regions incorrectly joins texts in wrong order.Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.16.16...0.16.17
Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a
np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.Update `unstructured-inference` to 0.8.6 in requirements which removed layoutparser dependency libs
unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libspdfminer-six to 20240706unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libspdfminer-six to 20240706Fix an issue with multiple values for `infer_table_structure` when paritioning email with image attachements the kwarg calls into partition to partiti
infer_table_structure when paritioning email with image attachements the kwarg calls into partition to partition the image already contains infer_table_structure. Now partition function checks if the kwarg has infer_table_structure alreadyAdd character-level filtering for tesseract output. It is controllable via TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.
TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.Prepare auto-partitioning for pluggable partitioners. Move toward a uniform partitioner call signature so a custom or override partitioner can be regi
application/vnd.ms-excel was incorrectly identified as an XLS file.Title elements.Title.Enhance quote standardization tests with additional Unicode scenarios
Table element was always segregated into its own pre-chunk such that the Table appeared alone in a chunk or was split into multiple TableChunk elements, but never combined with Text-subtype elements. Allow table elements to be combined with other elements in the same chunk when space allows.element.text. Previously .metadata.text_as_html was also considered and since it is always longer that the text (due to HTML tag overhead) it was the effective length criterion. Remove text-as-html from the length calculation such that text-length is the sole criterion for sizing a chunk.Fix original file doctype detection from cct converted file paths for metrics calculation.
chore: fix CHANGELOG formatting by @cragwolfe in https://github.com/Unstructured-IO/unstructured/pull/3800
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.16.8...0.16.9
Metrics: Weighted table average is optional
Add image_alt_mode to partition_html Adds an image_alt_mode parameter to partition_html() to control how alt text is extracted from images in HTML doc
image_alt_mode parameter to partition_html() to control how alt text is extracted from images in HTML documents for html_parser_version=v2 . The parameter can be set to to_text to extract alt text as text from <img> html tagsEvery ` ` tag is considered to be ontology.Table: Added special handling for tables in HTML partitioning. This change is made to improve the accuracy
<table> tag is considered to be ontology.Table: Added special handling for tables in HTML partitioning. This change is made to improve the accuracy of table extraction from HTML documents.UncategorizedText when the HTML tag is predicted correctly but has no class assigned.text_as_html metadata is combined across all elements in CompositeElement when chunking HTML output.<table> tag is considered to be ontology.Table Added special handling for tables in HTML partitioning (html_parser_version=v2. This change is made to improve the accuracy of table extraction from HTML documents.html_parser_version=v2 to ontology each defined HTML in the Ontology has assigned default ontology class. This way it is possible to assign ontology class instead of UncategorizedText when the HTML tag is predicted correctly without class assigned classtext_as_html metadata is combined across all elements in CompositeElement when chunking HTML outputAdd max recursion limit and fix to_text() method by @plutasnyy in https://github.com/Unstructured-IO/unstructured/pull/3773
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.16.4...0.16.5
`value` attribute in ` ` element is parsed to `OntologyElement.text` in ontology
value attribute in <input/> element is parsed to OntologyElement.text in ontologyid and class attributes removed from Table subtags in HTML partitioningto_html and newly introduced to_text in OntologyElementpartition_pdf() function now supports link extraction when using the hi_res strategy, allowing users to extract hyperlinks from PDF documents more effectively.V2 elements without first parent ID can be parsed
Whitespace-invariant CCT distance metric. CCT Levenshtein distance for strings is by default computed with standardized whitespaces.
Bump `unstructured-inference` to 0.7.39 and upgrade other dependencies
unstructured-inference to 0.7.39 and upgrade other dependenciespdfminer_processing.py to nearest machine precision. This can help reduce underterministic behavior from machine precision that affects which bounding boxes to combine.partition_via_api function. Expose retry-mechanism related parameters in the partition_via_api function to allow users to configure the retry behavior of the API requests.partition.email module and tests. Use modern Python stdlib email module interface to parse email messages and attachments. This change shortens and simplifies the code, and makes it more robust and maintainable. Several historical problems were remedied in the process..metadata.text_as_html for DOCX tables was "bloated" with whitespace and noise elements introduced by tabulate that produced over-chunking and lower "semantic density" of elements. Reduce HTML to minimum character count while preserving all text.filetype was incorrectly identified as a MSG file..metadata.text_as_html for DOCX tables was "bloated" with whitespace and noise elements introduced by pandas that produced over-chunking and lower "semantic density" of elements. Reduce HTML to minimum character count while preserving all text..metadata.text_as_html for CSV tables was "bloated" with whitespace and noise elements introduced by pandas that produced over-chunking and lower "semantic density" of elements. Reduce HTML to minimum character count while preserving all text..metadata.text_as_html for PPTX tables was "bloated" with whitespace and noise elements introduced by tabulate that produced over-chunking and lower "semantic density" of elements. Reduce HTML to minimum character count while preserving all text and structure.Remove ingest implementation. The deprecated ingest functionality has been removed, as it is now maintained in the separate unstructured-ingest reposi…
requirements/ingest directory with a new ingest.txt extra for installing the unstructured-ingest library.unstructured.ingest submodule.OCRAgentGoogleVision. Introduces an optional language parameter in the OCRAgentGoogleVision constructor to serve as a language hint for document_text_detection. This ensures compatibility with the OCRAgent's get_instance method and resolves errors when parsing PDFs with Google Cloud Vision as the OCR agent.Add (but do not install) a new post-partitioning decorator to handle metadata added for all file-types, like `.filename`, `.filetype` and `.languages`
.filename, .filetype and .languages. This will be installed in a closely following PR to replace the four currently being used for this purpose.partition_via_api. Make a minor syntax change to ensure forward compatibility with the upcoming 0.26.0 Python SDK.date_from_file_object parameter. As part of simplifying partitioning parameter set, remove date_from_file_object parameter. A file object does not have a last-modified date attribute so can never give a useful value. When a file-object is used as the document source (such as in Unstructured API) the last-modified date must come from the metadata_last_modified argument.KeyError when mapping parent ids to hash ids. Occasionally the input elements into assign_and_map_hash_ids can contain duplicated element instances, which lead to error when mapping parent id.@apply_metadata() decorator and only decorate the principal partitioner (CSV and DOCX in this case); remove decoration from delegating partitioners.@apply_metadata() decorator and only decorate the principal partitioner; remove decoration from delegating partitioners.@apply_metadata() decorator and only decorate the principal partitioner (HTML in this case); remove decoration from delegating partitioners.min_partition and max_partition parameters were an initial rough implementation of chunking but now interfere with chunking and are unused. Remove those parameters from partition_text() and partition_email().@apply_metadata() decorator operating on partitioners they delegate to (TXT, HTML, and all others for attachments) and remove direct decoration from EML and MSG.Improve `pdfminer` image cleanup process. Optimized the removal of duplicated pdfminer images by performing the cleanup before merging elements, rathe
pdfminer image cleanup process. Optimized the removal of duplicated pdfminer images by performing the cleanup before merging elements, rather than after. This improvement reduces execution time and enhances overall processing speed of PDF documents.numpy.float32 for coordinates and remove intermediate variables to reduce memory usage when computing intersection areasarm64 image build arm64 builds are now fixed and will be available against starting with the 0.15.13 release.file_utils.experimental and file_utils.metadata was removed. These functions were never published in the documentation, but if a client dug these out and used them this removal could break client code.pdfminer image cleanup process. Optimized the removal of duplicated pdfminer images by performing the cleanup before merging elements, rather than after. This improvement reduces execution time and enhances overall processing speed of PDF documents.numpy.float32 for coordinates and remove intermediate variables to reduce memory usage when computing intersection areasarm64 image build arm64 builds are now fixed and will be available against starting with the 0.15.13 release.Improve `pdfminer` element processing Implemented splitting of pdfminer elements (groups of text chunks) into smaller bounding boxes (text lines). Thi
pdfminer element processing Implemented splitting of pdfminer elements (groups of text chunks) into smaller bounding boxes (text lines). This prevents loss of information from the object detection model and facilitates more effective removal of duplicated pdfminer text.pdfminer element processing Implemented splitting of pdfminer elements (groups of text chunks) into smaller bounding boxes (text lines). This prevents loss of information from the object detection model and facilitates more effective removal of duplicated pdfminer text.Enhance `pdfminer` element cleanup Expand removal of pdfminer elements to include those inside all non-pdfminer elements, not just tables.
pdfminer element cleanup Expand removal of pdfminer elements to include those inside all non-pdfminer elements, not just tables.analysis of the partition_pdf function is set to True, the layout for Object Detection, Pdfminer Extraction, OCR and final layouts will be dumped as json files. The drawers now accept dict (dump) objects instead of internal classes instances.numpy operations to compute IOU and sub-region membership instead of using simply loop. This improves the speed of deduplicating elements for pages with a lot of elements.Add support for encoding parameter in partition_csv
detect_filetype to check the content of OLE files to more reliable differentiate DOC, PPT, XLS, and MSG files. As part of this, the "msg" extra was removed because the python-oxmsg package is now a base dependency.NamedTemporaryFile(..., delete=False) and/or uses of file.name of NamedTemporaryFiles have been replaced with TemporaryFileDirectory to avoid a known issue: https://docs.python.org/3/library/tempfile.html#tempfile.NamedTemporaryFileBump unstructured.paddleocr to 2.8.1.0.
pillow-heif with pi-heif. Replaces pillow-heif with pi-heif due to more permissive licensing on the wheel for pi-heif..metadata.text_as_html for DOCX tables was "bloated" with whitespace and noise elements introduced by tabulate that produced over-chunking and lower "semantic density" of elements. Reduce HTML to minimum character count without preserving all text.filetype was incorrectly identified as a MSG file.Fix NLTK data download path to prevent nested directories. Resolved an issue where a nested "nltk_data" directory was created within the parent "nltk_
Bump to NLTK 3.9.x Bumps to the latest nltk version to resolve CVE.
nltk version to resolve CVE.ingest-test-fixture-update-pr to resolve NLTK model download errors.TableChunk splits. When a Table element is divided during chunking to fit the chunking window, TableChunk.text corresponds exactly with the table text in TableChunk.metadata.text_as_html, .text_as_html is always parseable HTML, and the table is split on even row boundaries whenever possible.Revert to using `unstructured.pytesseract` fork. Due to the unavailability of some recent release versions of pytesseract on PyPI, the project now use
unstructured.pytesseract fork. Due to the unavailability of some recent release versions of pytesseract on PyPI, the project now uses the unstructured.pytesseract fork to ensure stability and continued support.libreoffice verson in image. Bumps the libreoffice version to 25.2.5.2 to address CVEs.nltk==3.8.2 on PyPI, the NLTK dependency has been downgraded to <3.8.2. This change ensures continued functionality and compatibility.Remove the custom index URL from `extra-paddleocr.in` to resolve the error in the `setup.py` configuration.
extra-paddleocr.in to resolve the error in the setup.py configuration.Mark ingest as deprecated Begin sunset of ingest code in this repo as it's been moved to a dedicated repo.
pdfminer embedded image extraction to exclude text elements and produce more accurate bounding boxes. This results in cleaner, more precise element extraction in pdf partitioning.Recipient elements are generated for cc and bcc when include_headers=True for email partitioning.pdf_hi_res_max_pages argument for partitioning, which allows rejecting PDF files that exceed this page number limit, when the high_res strategy is chosen. By default, it will allow parsing PDF files with an unlimited number of pages.HuggingFaceEmbeddingEncoder to use HuggingFaceEmbeddings from langchain_huggingface package instead of the deprecated version from langchain-community. This resolves the deprecation warning and ensures compatibility with future versions of langchain.OpenAIEmbeddingEncoder to use OpenAIEmbeddings from langchain-openai package instead of the deprecated version from langchain-community. This resolves the deprecation warning and ensures compatibility with future versions of langchain.detect_filetype() no longer silently falls back to detecting a file-type based on the extension when no file exists at the path provided. Instead FileNotFoundError is raised. This provides consistent user notification of a mis-typed path rather than an unpredictable exception from a file-type specific partitioner when the file cannot be opened.partition() as a file-path was identified as TXT and partitioned using partition_text(). EML files specified by path are now identified and processed correctly, including processing any attachments.partition() would raise when gzip compression was used for transport by the server.partition() with a swapped MS-Office content_type would cause the file-type to be misidentified. A DOCX, PPTX, or XLSX MIME-type received by partition() is now checked for accuracy and corrected if the file is for a different MS-Office 2007+ type.Improve text clearing process in email partitioning. Updated the email partitioner to remove both =\n and =\r\n characters during the clearing process
=\n and =\r\n characters during the clearing process. Previously, only =\n characters were removed.<p>, <div>) nested inside a phrasing element (e.g. <strong> or <cite>). Instead it breaks the phrasing run (and therefore element) at the block-item start and begins a new phrasing run after the block-item. This is consistent with how the browser determines element boundaries in this situation.partition_pdf(). Extend language specification capability to PaddleOCR in addition to TesseractOCR. Users can now specify OCR languages for both OCR engines when using partition_pdf().nltk binaries are downloaded. Work around a quirk in the Windows implementation of tempfile.NamedTemporaryFile where accessing the temporary file by name raises PermissionError.Remove NLTK download Removes nltk.download in favor of downloading from an S3 bucket we host to mitigate CVE-2024-39705
.doc files are now supported in the arm64 image.. libreoffice24 is added to the arm64 image, meaning .doc files are now supported. We have follow on work planned to investigate adding .ppt support for arm64 as well.nltk.download in favor of downloading from an S3 bucket we host to mitigate CVE-2024-39705.doc files are now supported in the arm64 image.. libreoffice24 is added to the arm64 image, meaning .doc files are now supported. We have follow on work planned to investigate adding .ppt support for arm64 as well.Add Object Detection Metrics to CI Add object detection metrics (average precision, precision, recall and f1-score) implementations.
nltk.download in favor of downloading from an S3 bucket we host to mitigate CVE-2024-39705Added visualization and OD model result dump for PDF In PDF hi_res strategy the analysis parameter can be used to visualize the result of the OD model
hi_res strategy the analysis parameter can be used to visualize the result of the OD model and dump the result to a file. Additionally, the visualization of bounding boxes of each layout source is rendered and saved for each page.partition_docx() distinguishes "file not found" from "not a ZIP archive" error. partition_docx() now provides different error messages for "file not found" and "file is not a ZIP archive (and therefore not a DOCX file)". This aids diagnosis since these two conditions generally point in different directions as to the cause and fix.soffice processes could be attempted Add a wait mechanism in convert_office_doc so that the function first checks if another soffice is running already: if yes wait till the other process finishes or till the wait timeout before spawning a subprocess to run sofficepartition() now forwards strategy arg to partition_docx(), partition_pptx(), and their brokering partitioners for DOC, ODT, and PPT formats. A strategy argument passed to partition() (or the default value "auto" assigned by partition()) is now forwarded to partition_docx(), partition_pptx(), and their brokering partitioners when those filetypes are detected.Move arm64 image to wolfi-base The arm64 image now runs on wolfi-base. The arm64 build for wolfi-base does not yet include libreoffce, and so arm64 do
arm64 image now runs on wolfi-base. The arm64 build for wolfi-base does not yet include libreoffce, and so arm64 does not currently support processing .doc, .ppt, or .xls file. If you need to process those files on arm64, use the legacy rockylinux image.Bump unstructured-inference==0.7.36 Fix ValueError when converting cells to html.
partition() now forwards strategy arg to partition_docx(), partition_ppt(), and partition_pptx(). A strategy argument passed to partition() (or the default value "auto" assigned by partition()) is now forwarded to partition_docx(), partition_ppt(), and partition_pptx() when those filetypes are detected.
Fix missing sensitive field markers for embedders
Pull from `wolfi-base` image. The amd64 image now pulls from the unstructured wolfi-base image to avoid duplication of dependency setup steps.
wolfi-base image. The amd64 image now pulls from the unstructured wolfi-base image to avoid duplication of dependency setup steps.extract_image_block_types and starting_page_number.wolfi-base image. The amd64 image now pulls from the unstructured wolfi-base image to avoid duplication of dependency setup steps.extract_image_block_types and starting_page_number.Remove deprecated `overwrite_schema` kwarg from Delta Table connector.. The overwrite_schema kwarg is deprecated in deltalake>=0.18.0. schema_mode= sh…
overwrite_schema kwarg from Delta Table connector.. The overwrite_schema kwarg is deprecated in deltalake>=0.18.0. schema_mode= should be used now instead. schema_mode="overwrite" is equivalent to overwrite_schema=True and schema_mode="merge" is equivalent to overwrite_schema="False". schema_mode defaults to None. You can also now specify engine, which defaults to "pyarrow". You need to specify enginer="rust" to use "schema_mode".partition_via_apiFiltering for tar extraction Adds tar filtering to the compression module for connectors to avoid decompression malicious content in .tar.gz files. Th
.tar.gz files. This was added to the Python tarfile lib in Python 3.12. The change only applies when using Python 3.12 and above.python-oxmsg for partition_msg(). Outlook MSG emails are now partitioned using the python-oxmsg package which resolves some shortcomings of the prior MSG parser.partition_msg() is now able to parse non-unicode Outlook MSG emails.partition_msg() is now able to extract attachments without corruption.Move logger error to debug level when PDFminer fails to extract text which includes error message for Invalid dictionary construct.
UnstructuredTableTransformerModel When a table is not recognized, the element.metadata.text_as_html attribute is set to an empty string.python-docx Pinned python-docx version to ensure a particular method unstructured uses is included.Add backward compatibility for the deprecated pdf_infer_table_structure parameter.
category field from Text class to Element class.partition_docx() now supports pluggable picture sub-partitioners. A subpartitioner that accepts a DOCX Paragraph and generates elements is now supported. This allows adding a custom sub-partitioner that extracts images and applies OCR or summarization for the image.partition_pdf() to keep spaces in the text. The control character \t is now replaced with a space instead of being removed when merging inferred elements with embedded elements.resolve_entities=False for XML parsing with lxml
to avoid text being dynamically injected into the XML document.form_extraction_skip_tables argument to the partition_pdf_or_image call.
to avoid text being dynamically injected into the XML document.table_as_cells output by default to reduce overhead in partition; now table_as_cells is only produced when the env EXTACT_TABLE_AS_CELLS is truedocument_to_element_list for handling HTMLDocument Use getattr(element, "type", "") to get the type attribute of an element when it exists. This is more explicit way to handle the special case for HTML documents and prevents other types of attribute error from being silenced by the try blockBump unstructured-inference==0.7.33.
pinecone connector.Nothing published for this version
Turn table extraction for PDFs and images off by default. Reverting the default behavior for table extraction to "off" for PDFs and images. A number o
partition_pdf(). Skip element sorting when determining whether embedded text can be extracted.partition_docx(). Behavior of future enhancements may be sensitive the partitioning strategy. Add this parameter so partition_docx() is aware of the requested strategy.NotImplementedError.paragraph_grouper can be set to False, but the type hint did not not reflect this previously.links is extracted during partitioning and is not needed as a paramter in partition_pdf.partition_csv() would raise on CSV files with very long lines.partition_doc(). Remove temporary file created but not removed when file argument is passed to partition_doc().SyntaxError or SyntaxWarning on regex patterns. Change regex patterns to raw strings to avoid these warnings/errors in Python 3.11+.partition_odt(). Remove temporary file created but not removed when file argument is passed to partition_odt().Remove `page_number` metadata fields for HTML partition until we have a better strategy to decide page counting.
page_number metadata fields for HTML partition until we have a better strategy to decide page counting.cid characters in embedded text extracted by pdfminer.partition_docx() handles short table rows. The DOCX format allows a table row to start late and/or end early, meaning cells at the beginning or end of a row can be omitted. While there are legitimate uses for this capability, using it in practice is relatively rare. However, it can happen unintentionally when adjusting cell borders with the mouse. Accommodate this case and generate accurate .text and .metadata.text_as_html for these tables.ValueError: Invalid file (FileType.UNK) when parsing Content-Type header with charset directive URL response Content-Type headers are now parsed accor
KeyError raised when updating parent_id In the past, combining ListItem elements could result in reusing the same memory location which then led to un
ListItem elements could result in reusing the same memory location which then led to unexpected side effects when updating element IDs.Unique and deterministic hash IDs for elements Element IDs produced by any partitioning function are now deterministic and unique at the document leve
basic and by_title. Remote chunking
options via the API are now accessible.UnstructuredTableTransformerModel is able to return predicted table in cells formatPDF_ANNOTATION_THRESHOLD environment variable to control the capture of embedded links in partition_pdf() for fast strategy.Remove duplicate image elements. Remove image elements identified by PDFMiner that have similar bounding boxes and the same text.
start_index in html links extractionstrategy arg value to _PptxPartitionerOptions. This makes this paritioning option available for sub-partitioners to come that may optionally use inference or other expensive operations to improve the partitioning.starting_page_number parameter to partitioning functions It applies to those partitioners which support page_number in element's metadata: PDF, TIFF, XLSX, DOC, DOCX, PPT, PPTX.unique_element_ids continues to be False by default, utilizing text hashes.<b> tags in HTML Now partition_html() can extract text from <b> tags inside container tags (like <div>, <pre>).Brings back missing word list files that caused partition failures in 0.13.1.
partition failures in 0.13.1.Drop constraint on pydantic, supporting later versions All dependencies has pydantic pinned at an old version. This explicit pin was removed, allowing
ElementTypes to extend future element typespartition_html() swallowing some paragraphs. The partition_html() only considers elements with limited depth to avoid becoming the text representation of a giant div. This fix increases the limit value.Add `.metadata.is_continuation` to text-split chunks. .metadata.is_continuation=True is added to second-and-later chunks formed by text-splitting an o
.metadata.is_continuation to text-split chunks. .metadata.is_continuation=True is added to second-and-later chunks formed by text-splitting an oversized Table element but not to their counterpart Text element splits. Add this indicator for CompositeElement to allow text-split continuation chunks to be identified for downstream processes that may wish to skip intentionally redundant metadata values in continuation chunks.compound_structure_acc metric to table eval. Add a new property to unstructured.metrics.table_eval.TableEvaluation: composite_structure_acc, which is computed from the element level row and column index and content accuracy scores.metadata.orig_elements to chunks. .metadata.orig_elements: list[Element] is added to chunks during the chunking process (when requested) to allow access to information from the elements each chunk was formed from. This is useful for example to recover metadata fields that cannot be consolidated to a single value for a chunk, like page_number, coordinates, and image_base64.--include_orig_elements option to Ingest CLI. By default, when chunking, the original elements used to form each chunk are added to chunk.metadata.orig_elements for each chunk. * The include_orig_elements parameter allows the user to turn off this behavior to produce a smaller payload when they don't need this metadata..metadata.orig_elements for each chunk. This behavior allows the text and metadata of the elements combined to make each chunk to be accessed. This can be important for example to recover metadata such as .coordinates that cannot be consolidated across elements and so is dropped from chunks. This option is controlled by the include_orig_elements parameter to partition_*() or to the chunking functions. This option defaults to True so original-elements are preserved by default. This behavior is not yet supported via the REST APIs or SDKs but will be in a closely subsequent PR to other unstructured repositories. The original elements will also not serialize or deserialize yet; this will also be added in a closely subsequent PR.clean_pdfminer_inner_elements() to remove only pdfminer (embedded) elements merged with inferred elements. Previously, some embedded elements were removed even if they were not merged with inferred elements. Now, only embedded elements that are already merged with inferred elements are removed.skip_infer_table_types parameter and reflect these changes in documentation.Improve ability to capture embedded links in `partition_pdf()` for `fast` strategy Previously, a threshold value that affects the capture of embedded
partition_pdf() for fast strategy Previously, a threshold value that affects the capture of embedded links was set to a fixed value by default. This allows users to specify the threshold value for better capturing.add_chunking_strategy decorator to dispatch by name. Add chunk() function to be used by the add_chunking_strategy decorator to dispatch chunking call based on a chunking-strategy name (that can be dynamic at runtime). This decouples chunking dispatch from only those chunkers known at "compile" time and enables runtime registration of custom chunkers..name not a local file path. When partitioning a file using the file= argument, and file is a file-like object (e.g. io.BytesIO) having a .name attribute, and the value of file.name is not a valid path to a file present on the local filesystem, FileNotFoundError is raised. This prevents use of the file.name attribute for downstream purposes to, for example, describe the source of a document retrieved from a network location via HTTP.pandoc which does not support RTF files + instructions that will help resolve that issue.install-pandoc Makefile recipe into relevant stages of CI workflow, ensuring it is a version that supports RTF input files.-1.0 with np.nan and corrected rows filtering of files metrics basing on that.partition_pdf() for fast strategy Previously, a threshold value that affects the capture of embedded links was set to a fixed value by default. This allows users to specify the threshold value for better capturing.add_chunking_strategy decorator to dispatch by name. Add chunk() function to be used by the add_chunking_strategy decorator to dispatch chunking call based on a chunking-strategy name (that can be dynamic at runtime). This decouples chunking dispatch from only those chunkers known at "compile" time and enables runtime registration of custom chunkers.table_level_acc metric for table evaluation. table_level_acc now is an average of individual predicted table's accuracy. A predicted table's accuracy is defined as the sequence matching ratio between itself and its corresponding ground truth table..name not a local file path. When partitioning a file using the file= argument, and file is a file-like object (e.g. io.BytesIO) having a .name attribute, and the value of file.name is not a valid path to a file present on the local filesystem, FileNotFoundError is raised. This prevents use of the file.name attribute for downstream purposes to, for example, describe the source of a document retrieved from a network location via HTTP.pandoc which does not support RTF files + instructions that will help resolve that issue.install-pandoc Makefile recipe into relevant stages of CI workflow, ensuring it is a version that supports RTF input files.-1.0 with np.nan and corrected rows filtering of files metrics basing on that.Header and footer detection for fast strategy partition_pdf with fast strategy now detects elements that are in the top or bottom 5 percent of the pag
partition_pdf with fast strategy now
detects elements that are in the top or bottom 5 percent of the page as headers and footers.identify_overlapping_or_nesting_case and catch_overlapping_and_nested_bboxes functions.partition_via_api() Update partition_via_api() to convert all list type parameters to JSON formatted strings before calling the unstructured client SDK. This will support image block extraction via partition_via_api().check_connection in opensearch, databricks, postgres, azure connectorscheck_connection in opensearch, databricks, postgres, azure connectors **partition_xlsx() that dropped content. Algorithm for detecting "subtables" within a worksheet dropped table elements for certain patterns of populated cells such as when a trailing single-cell row appeared in a contiguous block of populated cells.Key Concepts page.OpenAiEmbeddingConfig to OpenAIEmbeddingConfig.@add_chunking_strategy decorator was missing from partition_json() such that pre-partitioned documents serialized to JSON did not chunk when a chunking-strategy was specified.date_from_file_object parameter to partition. If True and if file is provided via file parameter it will cause partition to infer last modified date from file's content. If False, last modified metadata will be None.partition_pdf with fast strategy now
detects elements that are in the top or bottom 5 percent of the page as headers and footers.identify_overlapping_or_nesting_case and catch_overlapping_and_nested_bboxes functions.partition_via_api() Update partition_via_api() to convert all list type parameters to JSON formatted strings before calling the unstructured client SDK. This will support image block extraction via partition_via_api().check_connection in opensearch, databricks, postgres, azure connectorscheck_connection in opensearch, databricks, postgres, azure connectorspartition_xlsx() that dropped content. Algorithm for detecting "subtables" within a worksheet dropped table elements for certain patterns of populated cells such as when a trailing single-cell row appeared in a contiguous block of populated cells.Key Concepts page.OpenAiEmbeddingConfig to OpenAIEmbeddingConfig.@add_chunking_strategy decorator was missing from partition_json() such that pre-partitioned documents serialized to JSON did not chunk when a chunking-strategy was specified.…value; this function will be deprecated in the future and the default model name will simply rely on unstructured-inference and will not consider os e…
black formatting The black library recently introduced a new major version that introduces new formatting conventions. This change brings code in the unstructured repo into compliance with the new conventions..p7s files partition_email can now process .p7s files. The signature for the signed message is extracted and added to metadata.partition_email now falls back to anoter valid content type if it's available.OCRAgent interface and specify it using OCR_AGENT environment variable.partition_pdf() not working when using chipper model with filelanguages and ocr_languages Users are regularly receiving errors on the API because they are defining ocr_languages or languages with additional quotationmarks, brackets, and similar mistakes. This update handles common incorrect arguments and raises an appropriate warning.hi_res_model_name now relies on unstructured-inference When no explicit hi_res_model_name is passed into partition or partition_pdf_or_image the default model is picked by unstructured-inference's settings or os env variable UNSTRUCTURED_HI_RES_MODEL_NAME; it now returns the same model name regardless of infer_table_structure's value; this function will be deprecated in the future and the default model name will simply rely on unstructured-inference and will not consider os env in a future release.Driver for MongoDB connector. Adds a driver with unstructured version information to the MongoDB connector.
unstructured version information to the
MongoDB connector.unstructured-ingest to write partitioned data to a Databricks Volumes storage service.check_connection. There was an error when trying to ls destination directory - it may not exist at the moment of connector creation. Now check_connection calls ls on bucket root and this method is called on initialize of destination connector.setup.py is currently pointing to the wrong location for the databricks-volumes extra requirements. This results in errors when trying to build the wheel for unstructured. This change updates to point to the correct path.Fix index error in table processing. Bumps the unstructured-inference version to address and index error that occurs on some tables in the table trans
unstructured-inference version to address and
index error that occurs on some tables in the table transformer object.Drop support for python3.8 All dependencies are now built off of the minimum version of python being 3.10
Add SaaS API User Guide. This documentation serves as a guide for Unstructured SaaS API users to register, receive an API key and URL, and manage your
overlap_all kwarg on partition functions.Add intra-chunk overlap capability. Implement overlap for split-chunks where text-splitting is used to divide an oversized chunk into two or more chun
overlap kwarg on partition functions.image_base64 and image_mime_type (if that is what the user specifies by some other param like pdf_extract_to_payload). This would allow the API to have parity with the library.partition_via_api Users that self host the api were not able to pass their custom url to partition_via_api.Update the layout analysis script. The previous script only supported annotating final elements. The updated script also supports annotating inferred
final elements. The updated script also supports annotating inferred and extracted elements.staging_for brickstext/plain and text/html.unstructured-ingest to write partitioned/embedded data to a Chroma vector database.Fix `partition_pdf()` and `partition_image()` importation issue. Reorganize pdf.py and image.py modules to be consistent with other types of document
partition_pdf() and partition_image() importation issue. Reorganize pdf.py and image.py modules to be consistent with other types of document import code.Refactor image extraction code. The image extraction code is moved from unstructured-inference to unstructured.
unstructured-inference to unstructured.unstructured-inference to unstructured.sensitive annotation for fields related to auth (i.e. passwords, tokens). Refactor all fsspec connectors to use explicit access configs rather than a generic dictionary.table-<pageN>-<tableN>.jpg. This filename is presented in the image_path metadata field for the Table element. The default would be to not do this.unstructured-ingest to write partitioned data from over 20 data sources (so far) to a Weaviate object collection.hi_res partitioning failure when pdfminer fails. Implemented logic to fall back to the "inferred_layout + OCR" if pdfminer fails in the hi_res strategy.tesseract can handle for ocr layout detectionpartition_csv now identifies the correct delimiter before the file is processed.hi_res occasionally pdfminer can fail to decode the text in an pdf file and return cid code as text. Now when this happens the text from OCR is used.Your coding agent can read these notes before it upgrades. Set up the MCP server →