NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2020 most downloaded on PyPI
A library that prepares raw documents for downstream ML tasks.
Last release 3 days ago
14 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
3 versions withdrawn
withdrawn after publishing
4 years old
233 releases · first in 2022
One column per quarter.
fast strategy for pdf now keeps element bounding box data
Add partition_csv for CSV files.
partition_csv for CSV files.Deprecate --s3-url in favor of --remote-url in CLI
--s3-url in favor of --remote-url in CLIfile_directory to metadatapage_name to metadata. Currently used for the sheet name in XLSX documents.--partition-strategy parameter to unstructured-ingest so that users can specify
partition strategy in CLI. For example, --partition-strategy fast.unstructured/file-utils/filetype.py to better utilise hashmap to return mime type.test_filetype.py.partition_xml for XML files.partition_xlsx for Microsoft Excel documents.hml filetype for partition as a variation of html filetype.pytesseract a function level import in partition_pdf so you can use the "fast"
or "hi_res" strategies if pytesseract is not installed. Also adds the
required_dependencies decorator for the "hi_res" and "ocr_only" strategies.filename is tracked in metadata for docx tables.Adds an "auto" strategy that chooses the partitioning strategy based on document characteristics and function kwargs. This is the new default strategy
"auto" strategy that chooses the partitioning strategy based on document
characteristics and function kwargs. This is the new default strategy for partition_pdf
and partition_image. Users can maintain existing behavior by explicitly setting
strategy="hi_res".get_date method to ElementMetadata for converting the datestring to a datetime object.filename attribute on ElementMetadata to remove the full filepath.partition_docx in docxfileutils/file_type check json and eml decode ignore errorpartition_email was updated to more flexibly handle deviations from the RFC-2822 standard.
The time in the metadata returns None if the time does not match RFC-2822 at all.Added support for SpooledTemporaryFile file argument.
Added an "ocr_only" strategy for partition_pdf. Refactored the strategy decision logic into its own module.
partition_pdf. Refactored the strategy decision
logic into its own module.Add an "ocr_only" strategy for partition_image.
partition_image.partition_multiple_via_api for partitioning multiple documents in a single REST
API call.stage_for_baseplate function to prepare outputs for ingestion into Baseplate.partition_odt for processing Open Office documents.partition_pdf fast strategy to group together text
in the same bounding box.Added logic to partition_pdf for detecting copy protected PDFs and falling back to the hi res strategy when necessary.
partition_pdf for detecting copy protected PDFs and falling back
to the hi res strategy when necessary.partition_via_api for partitioning documents through the hosted API.exceeds_cap_ratio handles empty (returns True instead of False)detect_filetype to properly detect JSONs when the MIME type is text/plain.Updated the table extraction parameter name to be more descriptive
Adds an ssl_verify kwarg to partition and partition_html to enable turning off SSL verification for HTTP requests. SSL verification is on by default.
ssl_verify kwarg to partition and partition_html to enable turning off
SSL verification for HTTP requests. SSL verification is on by default.partition_pdf and partition_image through
the ocr_language kwarg. ocr_language corresponds to the code for the language pack
in Tesseract. You will need to install the relevant Tesseract language pack to use a
given language.partition and partition_pdf..msg filesAllow headers to be passed into partition when url is used.
partition when url is used.bytes_string_to_string cleaning brick for bytes string output.exactly_one in partition_jsonNone in _read_xml._read_xml so that Markdown files with embedded HTML process correctly.partition_pdf and partition_text group broken paragraphs to avoid fragmented NarrativeText elements.Add OS mimetypes DB to docker image, mainly for unstructured-api compat.
partition_text to group together broken paragraphs.partition_rtf for processing rich text files.partition now accepts a url kwarg in addition to file and filename.replace_mime_encodings.partition_text to group together broken paragraphs.partition_rtf for processing rich text files.partition now accepts a url kwarg in addition to file and filename.replace_mime_encodings.Guard against null style attribute in docx document elements
Add sender, recipient, date, and subject to element metadata for emails
--download-only parameter to unstructured-ingestConvert file to str in helper split_by_paragraph for partition_text
split_by_paragraph for partition_textUpdate elements_to_json to return string when filename is not specified
elements_to_json to return string when filename is not specifiedelements_from_json may take a string instead of a filename with the text kwargdetect_filetype now does a final fallback to file extension.unstructured-ingest--max-docs parameter to unstructured-ingestpartition_msg for processing MSFT Outlook .msg files.convert_file_to_text now passes through the source_format and target_format kwargs.
Previously they were hard coded.text kwarg no longer raise an error if an empty
string is passed (and empty list of elements is returned instead).partition_json no longer fails if the input is an empty list.chunk_by_attention_window that caused the last word in segments to be cut-off
in some cases.stage_for_transformers now returns a list of elements, making it consistent with other
staging bricksRefactored codebase using exactly_one
exactly_onecontent_type and file_filename parameters to partition() to bypass file detection--flatten-metadata parameter to unstructured-ingest--fields-include parameter to unstructured-ingestFix problem with PDF partition (duplicated test)
contains_english_word(), used heavily in text processing, is 10x faster.--metadata-include and --metadata-exclude parameters to unstructured-ingestclean_non_ascii_chars to remove non-ascii characters from unicode stringpartition_pdf(..., strategy="fast")Start deprecation life cycle for unstructured-ingest --s3-url option, to be deprecated in favor of --remote-url.
FsspecConnector to easily integrate any existing fsspec filesystem as a connector.s3_connector.py to s3.py for readability and consistency with the
rest of the connectors.S3Connector relies on s3fs instead of on boto3, and it inherits
from FsspecConnector.UNSTRUCTURED_LANGUAGE_CHECKS environment variable to control whether or not language
specific checks like vocabulary and POS tagging are applied. Set to "true" for higher
resolution partitioning and "false" for faster processing.detect_filetype warning to include filename when provided.unstructured-ingest --s3-url option, to be deprecated in
favor of --remote-url.AzureBlobStorageConnector based on its fsspec implementation inheriting
from FsspecConnectorpartition_epub for partitioning e-books in EPUB3 format.message/rfc822 MIME type.auto.partition() can now load Unstructured ISD json documents.
auto.partition() can now load Unstructured ISD json documents.--wikipedia-auto-suggest argument to the ingest CLI to disable automatic redirection
to pages with similar names.encoding argument to the partition_(text/email/html) functions.unstructured-ingest now uses a default --download_dir of $HOME/.cache/unstructured/ingest rather than a "tmp-ingest-" dir in the working directory.
unstructured-ingest now uses a default --download_dir of $HOME/.cache/unstructured/ingest
rather than a "tmp-ingest-" dir in the working directory.setup_ubuntu.sh no longer fails in some contexts by interpreting
DEBIAN_FRONTEND=noninteractive as a commandunstructured-ingest no longer re-downloads files when --preserve-downloads
is used without --download-dir.Fixes an error causing JavaScript to appear in the output of partition_html sometimes.
partition_html sometimes.requires_dependencies decorator, including the error message
and how it was used, which had caused an error for unstructured-ingest --github-url ....Add requires_dependencies Python decorator to check dependencies are installed before instantiating a class or running a function
requires_dependencies Python decorator to check dependencies are installed before
instantiating a class or running a functionprocess_document file cleaning on failureNarrativeText
and FigureCaption elements to be represented as Text in HTML documents.Fallback to using file extensions for filetype detection if libmagic is not present
libmagic is not presentpartition_md partitioner.Added elements_to_json and elements_from_json for easier serialization/deserialization
elements_to_json and elements_from_json for easier serialization/deserializationconvert_to_dict, dict_to_elements and convert_to_csv are now aliases for functions
that use the ISD terminology.Automatically install nltk models in the tokenize module.
nltk models in the tokenize module.## 0.4.13 * Fixes unstructured-ingest cli.
Adds console_entrypoint for unstructured-ingest, other structure/doc updates related to ingest.
Adds partition_doc for partitioning Word documents in .doc format. Requires libreoffice.
partition_doc for partitioning Word documents in .doc format. Requires libreoffice.partition_ppt for partitioning PowerPoint documents in .ppt format. Requires libreoffice.Fixes ElementMetadata so that it's JSON serializable when the filename is a Path object.
ElementMetadata so that it's JSON serializable when the filename is a Path object.Added ingest modules and s3 connector
url=None for partition_pdf and partition_imageUNSTRUCTURED_LANGUAGE env var to "".Element objects now track metadataurl=None for partition_pdf and partition_imageUNSTRUCTURED_LANGUAGE env var to "".Element objects now track metadataModified XML and HTML parsers not to load comments.
Added the ability to pull an HTML document from a url in partition_html.
partition_html.partition for .pptx, .pdf, images, and .html files.to_dict method to document elements.replace_unicode_quotes.Loosen the default cap threshold to 0.5.
0.5.UNSTRUCTURED_NARRATIVE_TEXT_CAP_THRESHOLD environment variable for controlling
the cap ratio threshold.Text for HTML and plain text documents.Body Text styles no longer default to NarrativeText for Word documents. The style information
is insufficient to determine that the text is narrative.Address element for capturing elements that only contain an address.UserWarning when detectron is called.UNSTRUCTURED_TITLE_MAX_WORD_LENGTH
environment variable for controlling the max number of words in a title.partition_pptx to order the elements on the pageUpdated partition_pdf and partition_image to return unstructured Element objects
partition_pdf and partition_image to return unstructured Element objectscoordinates attribute to document objectsFigureCaption and CheckBox document elementsLayoutElement objectspartition_pptx for partitioning PowerPoint documentsAdds requests as a base dependency
requests as a base dependencyexceeds_cap_ratio so the function doesn't break with empty text_parse_received_data.detect_filetype to properly handle .doc, .xls, and .ppt.Added partition_image to process documents in an image format.
partition_image to process documents in an image format.partition_email with attachments for text/htmlAdded support for text files in the partition function
partition functionopencv-python for easier installation on LinuxAdded generic partition brick that detects the file type and routes a file to the appropriate partitioning brick.
partition brick that detects the file type and routes a file to the appropriate
partitioning brick.partition_html and partition_eml to support file-like objects in 'rb' mode.clean_ordered_bullets.extract_ordered_bullets.clean_ordered_bullets.extract_ordered_bullets.partition_docx for pre-processing Word Documents.parse_received_data and partition_headerpartition_textextract_ip_address, extract_ip_address_name, extract_mapi_id, extract_datetimetzImage element and function to find embedded images find_embedded_imagesget_directory_file_info for summarizing information about source documentsAdd support for local inference
partition_html that allows for processing div tags that have both text and child
elements.docx, .xlsx, and .jpg files.extract_attachment_info that extracts and decodes the attachment
of an email.Elements to a pandas dataframe.partition_email## 0.3.4 * Python-3.7 compat
Removes BasicConfig from logger configuration
partition_email partitioning brickreplace_mime_encodings cleaning brickspartition_email partitioning brickreplace_mime_encodings cleaning bricksEmailElement data structure to store email documentsAdded translate_text brick for translating text between languages
translate_text brick for translating text between languagesapply method to make it easier to apply cleaners to elementsAdded \_\_init.py\_\_ to partition
partitionImplement staging brick for Argilla. Converts lists of Text elements to argilla dataset classes.
Text elements to argilla dataset classes.replace_unicode_quotes brickpartition_html for partitioning HTML documents.Nothing published for this version
Update python requirement to >=3.7
Add an alternative way of importing Final to support google colab
Final to support google colabAdd cleaning bricks for removing prefixes and postfixes
### 0.2.2 * Add staging brick for Datasaur
Added brick to convert an ISD dictionary to a list of elements
PDFDocument to use the from_file methodtransformers.Initial release of unstructured
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →