NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2528 most downloaded on PyPI
A library that prepares raw documents for downstream ML tasks.
Last release 7 days ago
27 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
3 versions withdrawn
withdrawn after publishing
4 years old
236 releases · first in 2022
One column per quarter.
Update the layout analysis script. The previous script only supported annotating final elements. The updated script also supports annotating inferred
final elements. The updated script also supports annotating inferred and extracted elements.staging_for brickstext/plain and text/html.unstructured-ingest to write partitioned/embedded data to a Chroma vector database.Fix `partition_pdf()` and `partition_image()` importation issue. Reorganize pdf.py and image.py modules to be consistent with other types of document
partition_pdf() and partition_image() importation issue. Reorganize pdf.py and image.py modules to be consistent with other types of document import code.Refactor image extraction code. The image extraction code is moved from unstructured-inference to unstructured.
unstructured-inference to unstructured.unstructured-inference to unstructured.sensitive annotation for fields related to auth (i.e. passwords, tokens). Refactor all fsspec connectors to use explicit access configs rather than a generic dictionary.table-<pageN>-<tableN>.jpg. This filename is presented in the image_path metadata field for the Table element. The default would be to not do this.unstructured-ingest to write partitioned data from over 20 data sources (so far) to a Weaviate object collection.hi_res partitioning failure when pdfminer fails. Implemented logic to fall back to the "inferred_layout + OCR" if pdfminer fails in the hi_res strategy.tesseract can handle for ocr layout detectionpartition_csv now identifies the correct delimiter before the file is processed.hi_res occasionally pdfminer can fail to decode the text in an pdf file and return cid code as text. Now when this happens the text from OCR is used.Updated Documentation: (i) Added examples, and (ii) API Documentation, including Usage, SDKs, Azure Marketplace, and parameters and validation errors.
Use `pikepdf` to repair invalid PDF structure for PDFminer when we see error PSSyntaxError when PDFminer opens the document and creates the PDFminer p
Use pikepdf to repair invalid PDF structure for PDFminer when we see error PSSyntaxError when PDFminer opens the document and creates the PDFminer pages object or processes a single PDF page.
Batch Source Connector support For instances where it is more optimal to read content from a source connector in batches, a new batch ingest doc is added which created multiple ingest docs after reading them in in batches per process.
<style> tags in HTML. <style> tags containing CSS in invalid positions previously contributed to element text. Do not consider text node of a <style> element as textual content.Header and Footer document elements.Add a class for the strategy constants. Add a class PartitionStrategy for the strategy constants and use the constants to replace strategy strings.
PartitionStrategy for the strategy constants and use the constants to replace strategy strings.DEFAULT_PADDLE_LANG before we have the language mapping for paddle.element.metadata.coefficient = 0.58. These fields will round-trip through JSON and can be accessed with dotted notation.TYPE_TO_TEXT_ELEMENT_MAP Updated Figure mapping from FigureCaption to Image.pdfminer, causing partition_pdf() to fail. We expect to be able to partition smoothly using an alternative strategy if text extraction doesn't work. Added exception handling to handle unexpected errors when extracting pdf text and to help determine pdf strategy.fast strategy fall back to ocr_only The fast strategy should not fall back to a more expensive strategy.languages in metadata when partitioning strategy=hi_res or fast User defined languages was previously used for text detection, but not included in the resulting element metadata for some strategies. languages will now be included in the metadata regardless of partition strategy for pdfs and images.KeyError: 'N' Certain pdfs were throwing this error when being opened by pdfminer. Added a wrapper function for pdfminer that allows these documents to be partitioned.Table chunks. Remedies repeated appearance of full .text_as_html on metadata of each TableChunk split from a Table element too large to fit in the chunking window.<thead> or <tfoot> element was emitted as a table element having no text and unparseable HTML in element.metadata.text_as_html. Do not emit empty tables to the element stream.element.metadata.text_as_html contains spurious <br> elements in invalid locations. The HTML generated for the text_as_html metadata for HTML tables contained <br> elements invalid locations like between <table> and <tr>. Change the HTML generator such that these do not appear.<thead> or <tfoot> element were not detected and the text in those cells was omitted from the table element text and .text_as_html. Detect table rows regardless of the semantic tag they may be nested in..text_as_html. tabulate inserts padding spaces to achieve visual alignment of columns in HTML tables it generates. Add our own HTML generator to do this simple job and omit that padding as well as newlines ("\n") used for human readability.output-dir/input-filename.jsonSupport nested DOCX tables. In DOCX, like HTML, a table cell can itself contain a table. In this case, create nested HTML tables to reflect that struc
check_connection() method which makes sure a valid connection can be established with the source/destination given any authentication credentials in a lightweight request.url. partition now accepts a new optional parameter request_timeout which if set will prevent any requests.get from hanging indefinitely and instead will raise a timeout error. This is useful when partitioning a url that may be slow to respond or may not respond at all._determine_pdf_auto_strategy returned hi_res strategy only if infer_table_structure was true. It now returns the hi_res strategy if either infer_table_structure or extract_images_in_pdf is true.0. A logical check is now added to avoid such error.ocr_only Tables that contain only numbers are returned as floats in a pandas.DataFrame when the image is converted from .image_to_data(). An AttributeError was raised downstream when trying to .strip() the floats.w:lastRenderedPageBreak elements present in the document XML. Page breaks are NOT reliably indicated by "hard" page-breaks inserted by the author and when present are redundant to a w:lastRenderedPageBreak element so cause over-counting if used. Use rendered page-breaks only.Add include_header argument for partition_csv and partition_tsv Now supports retaining header rows in CSV and TSV documents element partitioning.
SourceConnectionNetworkError custom error, which triggers the retry logic, if enabled, in the ingest pipeline.additional_partition_args arg was added to allow users to pass in any other arguments that should be added when calling partition(). This helps keep any changes to the input parameters of the partition() exposed in the CLI.partition_text to prevent empty elements Adds a check to filter out empty bullets.ocr_languages with values for languages Some API users ran into an issue with sending languages params because the API defaulted to also using an empty string for ocr_languages. This update handles situations where languages is defined and ocr_languages is an empty string.annots that resolved out as None. A logical check added to avoid such error.Estimating resolution as X leaded by invalid language parameters input. Proceed with defalut language eng when lang.py fails to find valid language code for tesseract, so that we don't pass an empty string to tesseract CLI and raise an exception in downstream.Add element type CI evaluation workflow Adds element type frequency evaluation metrics to the current ingest workflow to measure the performance of ea
yolox by default for table extraction when partitioning pdf/image yolox model provides higher recall of the table regions than the quantized version and it is now the default element detection model when infer_table_structure=True for partitioning pdf/image fileshi_res some elements where extracted using pdfminer too, so we removed pdfminer from the tables pipeline to avoid duplicated elements.unstructured-ingest to write to any of the following:
ocr_only strategy in partition_pdf() Adds the functionality to get accurate coordinate data when partitioning PDFs and Images with the ocr_only strategy.tables extension when instantiating the python-markdown object. Importance: This will allow users to extract structured data from tables in markdown documents.get_uris_from_annots function tried to access the dictionary value of a string instance variable. Assign None to the annotation variable if the instance type is not dictionary to avoid the erroneous attempt.yolox by default for table extraction when partitioning pdf/image yolox model provides higher recall of the table regions than the quantized version and it is now the default element detection model when infer_table_structure=True for partitioning pdf/image fileshi_res some elements where extracted using pdfminer too, so we removed pdfminer from the tables pipeline to avoid duplicated elements.unstructured-ingest to write to any of the following:
ocr_only strategy in partition_pdf() Adds the functionality to get accurate coordinate data when partitioning PDFs and Images with the ocr_only strategy.tables extension when instantiating the python-markdown object. Importance: This will allow users to extract structured data from tables in markdown documents.get_uris_from_annots function tried to access the dictionary value of a string instance variable. Assign None to the annotation variable if the instance type is not dictionary to avoid the erroneous attempt.Leverage dict to share content across ingest pipeline To share the ingest doc content across steps in the ingest pipeline, this was updated to use a m
ebooklib as a dependency ebooklib is licensed under AGPL3, which is incompatible with the Apache 2.0 license. Thus it is being removed.re_download to dictate if files should be forced to redownload rather than use what might already exist locally.Add CI evaluation workflow Adds evaluation metrics to the current ingest workflow to measure the performance of each file extracted as well as aggrega
overlapping_elements, overlapping_case, overlapping_percentage, largest_ngram_percentage, overlap_percentage_total, max_area, min_area, and total_area.typing-extensions as an explicit dependency This package is an implicit dependency, but the module is being imported directly in unstructured.documents.elements so the dependency should be explicit in case changes in other dependencies lead to typing-extensions being dropped as a dependency.extract_tables to unstructured-inference since it is now supported in unstructured instead Table extraction previously occurred in unstructured-inference, but that logic, except for the table model itself, is now a part of the unstructured library. Thus the parameter triggering table extraction is no longer passed to the unstructured-inference package. Also noted the table output regression for PDF files.skip_infer_table_types variable used in partition was not being passed down to specific file partitioners. Now you can utilize the skip_infer_table_types list variable when calling partition to specify the filetypes for which you want to skip table extraction, or the infer_table_structure boolean variable on the file specific partitioning function.max_characters.Duplicate CLI param check Given that many of the options associated with the Click based cli ingest commands are added dynamically from a number of co
Click based cli ingest commands are added dynamically from a number of configs, a check was incorporated to make sure there were no duplicate entries to prevent new configs from overwriting already added options.OCR_AGENT for OCRing the entire document.calculate_edit_distance. For easy function call, it is now a wrapper around the original function that calls edit_distance and return as "score".unstructured.embed.bedrock now provides a connector to use AWS bedrock's titan-embed-text model to generate embeddings for elements. This features requires valid AWS bedrock setup and an internet connectionto run.PDFResourceManager from pdfminer.converter which was causing an error for some users. We changed to import from the actual location of PDFResourceManager, which is pdfminer.pdfinterp.langdetect if the language was attempted to be detected on an empty string. Language detection is now skipped for empty strings.regex_metadata was used, where every element that contained a regex-match would start a new chunk.__init__.py file under the folder.partition_pdf get model_name=None In API usage the model_name value is None and the cast function in partition_pdf would return None and lead to attribution error. Now we use str function to explicit convert the content to string so it is guaranteed to have starts_with and other string functions as attributestbody tag HTML tables may sometimes just contain headers without body (tbody tag)max_characters.Improve natural reading order Some OCR elements with only spaces in the text have full-page width in the bounding box, which causes the xycut sorting
OCR elements with only spaces in the text have full-page width in the bounding box, which causes the xycut sorting to not work as expected. Now the logic to parse OCR results removes any elements with only spaces (more than one space).__init__.py file under the folder.Add functionality to limit precision when serializing to json Precision for points is limited to 1 decimal point if coordinates["system"] == "PixelSpa
points is limited to 1 decimal point if coordinates["system"] == "PixelSpace" (otherwise 2 decimal points?). Precision for detection_class_prob is limited to 5 decimal points.under_non_alpha_ratio dividing by zero Although this function guarded against a specific cause of division by zero, there were edge cases slipping through like strings with only whitespace. This update more generally prevents the function from performing a division by zero.bump `unstructured-inference` to `0.7.3` The updated version of unstructured-inference supports a new version of the Chipper model, as well as a clean
unstructured-inference to 0.7.3 The updated version of unstructured-inference supports a new version of the Chipper model, as well as a cleaner schema for its output classes. Support is included for new inference features such as hierarchy and ordering.--skip-infer-table-types parameter was added to map to the skip_infer_table_types partition argument. This gives more granular control to unstructured-ingest users, allowing them to specify the file types for which we should attempt table extraction.metadata.links, metadata.link_texts and metadata.link_urls for elements that contain a hyperlink that points to an external resource. So-called "jump" links pointing to document internal locations (such as those found in a table-of-contents "jumping" to a chapter or section) are excluded.Add elements_to_text as a staging helper function In order to get a single clean text output from unstructured for metric calculations, automate the process of extracting text from elements using this function.
Adds permissions(RBAC) data ingestion functionality for the Sharepoint connector. Problem: Role based access control is an important component in many data storage systems. Users may need to pass permissions (RBAC) data to downstream systems when ingesting data. Feature: Added permissions data ingestion functionality to the Sharepoint connector.
run method. This allows users to specify that connectors fetch embeddings without failure.__init__.py in order to make it discoverable.## 0.10.21 * Adds Scarf analytics.
Add document level language detection functionality. Adds the "auto" default for the languages param to all partitioners. The primary language present
langdetect package. Additional param detect_language_per_element is also added for partitioners that return multiple elements. Defaults to False.xy-cut sorting: Update shrink_bbox() to keep top left rather than center.max_characters=<n> argument to all element types in add_chunking_strategy decorator Previously this argument was only utilized in chunking Table elements and now applies to all partitioned elements if add_chunking_strategy decorator is utilized, further preparing the elements for downstream processing.bag_of_words and percent_missing_text functions In order to count the word frequencies in two input texts and calculate the percentage of text missing relative to the source document.edit_distance calculation metrics In order to benchmark the cleaned, extracted text with unstructured, edit_distance (Levenshtein distance) is included.metadata module had several top level imports that were only used in and applicable to code related to specific document types, while there were many general-purpose functions. As a result, general-purpose functions couldn't be used without unnecessary dependencies being installed. Fix: moved 3rd party dependency top level imports to inside the functions in which they are used and applied a decorator to check that the dependency is installed and emit a helpful error message if not.Title elements from chipper get category_depth= None even when Headline and/or Subheadline elements are present in the same page. Fix: all Title elements with category_depth = None should be set to have a depth of 0 instead iff there are Headline and/or Subheadline element-types present. Importance: Title elements should be equivalent html H1 when nested headings are present; otherwise, category_depth metadata can result ambiguous within elements in a page.xy-cut ordering output to be more column friendly This results in the order of elements more closely reflecting natural reading order which benefits downstream applications. While element ordering from xy-cut is usually mostly correct when ordering multi-column documents, sometimes elements from a RHS column will appear before elements in a LHS column. Fix: add swapped xy-cut ordering by sorting by X coordinate first and then Y coordinate.GoToR which refers to pdf resources outside of its own was detected since no condition catches such case. The code is fixing the issue by initialize URI before any condition check.Adds XLSX document level language detection Enhancing on top of language detection functionality in previous release, we now support language detectio
.xlsx file type at Element level.unstructured-inference to 0.6.6 The updated version of unstructured-inference makes table extraction in hi_res mode configurable to fine tune table extraction performance; it also improves element detection by adding a deduplication post processing step in the hi_res partitioning of pdfs and images.add_chunking_strategy decorator to partition functions. In addition to combining elements under Title elements, user's can now specify the max_characters=<n> argument to chunk Table elements into TableChunk elements with text and text_as_html of length <n> characters. This means partitioned Table results are ready for use in downstream applications without any post processing.hi_res model for pdf/image partition to yolox Now partitioning pdf/image using hi_res strategy utilizes yolox_quantized model isntead of detectron2_onnx model. This new default model has better recall for tables and produces more detailed categories for elements.partition_xlsx now can reads subtable(s) within one .xlsx sheet, along with extracting other title and narrative texts. Importance: This enhance the power of .xlsx reading to not only one table per sheet, allowing user to capture more data tables from the file, if exists.partition_pdf when attempt to get bounding box from element experienced a reference before assignment error when the first object is not text extractable. Fix: Switched to a flag when the condition is met. Importance: Crucial to be able to partition with pdf.detection_class_prob appears in Element metadata Problem: when detection_class_prob appears in Element metadata, Elements will only be combined by chunk_by_title if they have the same detection_class_prob value (which is rare). This is unlikely a case we ever need to support and most often results in no chunking. Fix: detection_class_prob is included in the chunking list of metadata keys excluded for similarity comparison. Importance: This change allows chunk_by_title to operate as intended for documents which include detection_class_prob metadata in their Elements.Nothing published for this version
Fix: remove nonexistent fields before instantiating in ElementMetadata.from_json(). Importance: Crucial to avoid breaking changes when adding fields.
xy-cut sorting to preprocess bboxes, shrinking all bounding boxes by 90% along x and y axes (still centered around the same center point), which allows projection lines to be drawn where not possible before if layout bboxes overlapped.partition_xml to be faster and more memory efficient when partitioning large XML files The new behavior is to partition iteratively to prevent loading the entire XML tree into memory at once in most use cases.unstructured-ingest to write partitioned data from over 20 data sources (so far) to an Azure Cognitive Search index.langdetect package. Adds the document languages as ISO 639-3 codes to the element metadata. Implemented only for the partition_text function to start.links metadata in partition_pdf for fast strategy. Problem: PDF files contain rich information and hyperlink that Unstructured did not captured earlier. Feature: partition_pdf now can capture embedded links within the file along with its associated text and page number. Importance: Providing depth in extracted elements give user a better understanding and richer context of documents. This also enables user to map to other elements within the document if the hyperlink is refered internally.pip install unstructured[all-docs] it will now upgrade both unstructured and unstructured-inference. Importance: This will ensure that the inference library is always in sync with the unstructured library, otherwise users will be using outdated libraries which will likely lead to unintended behavior.__post_init__. Fix: Adds a try/catch when the IngestConnector runs get_ingest_docs such that the error is logged but all processable documents->IngestDocs are still instantiated and returned. Importance: Allows users to ingest SharePoint content even when some files with unsupported filetypes exist there.make html and installs library to suppress warnings.partition_via_api, the hosted api may return an element schema that's newer than the current unstructured. In this case, metadata fields were added which did not exist in the local ElementMetadata dataclass, and __init__() threw an error. Fix: remove nonexistent fields before instantiating in ElementMetadata.from_json(). Importance: Crucial to avoid breaking changes when adding fields.None Problem: Getting the jump_url from a nonexistent Discord channel fails. Fix: property jump_url is now retrieved within the same context as the messages from the channel. Importance: Avoids cascading issues when the connector fails to fetch information about a Discord channel.deltalake on Linux Problem: occasionally on Linux ingest can throw a SIGABTR when writing deltalake table even though the table was written correctly. Fix: put the writing function into a Process to ensure its execution to the fullest extent before returning to the main process. Importance: Improves stability of connectors using deltalakeAdds data source properties to Airtable, Confluence, Discord, Elasticsearch, Google Drive, and Wikipedia connectors These properties (date_created, da
languages param in any Tesseract-supported langcode or any ISO 639 standard language code.languages param in any Tesseract-supported langcode or any ISO 639 standard language code.langdetect package. Implemented only for the partition_text function to start.Adds `languages` as an input parameter and marks `ocr_languages` kwarg for deprecation in pdf, image, and auto partitioning functions. Previously, lan…
unstructured element categories so the consumer of the library would see many UncategorizedText elements. This fixes the issue, improving the granularity of the element categories outputs for better downstream processing and chunking. The mapping update is:
NarrativeTextNarrativeTextTitleNarrativeTextNarrativeTextTitle (with category_depth=1)Title (with category_depth=2)NarrativeTextpartition_pdf with fast strategy previously broke down some numbered list item lines as separate elements. This enhancement leverages the x,y coordinates and bbox sizes to help decide whether the following chunk of text is a continuation of the immediate previous detected ListItem element or not, and not detect it as its own non-ListItem element.UncategorizedText elements.Table Elements are now propery extracted.add_chunking_strategy decorator to partition functions. Previously, users were responsible for their own chunking after partitioning elements, often required for downstream applications. Now, individual elements may be combined into right-sized chunks where min and max character size may be specified if chunking_strategy=by_title. Relevant elements are grouped together for better downstream results. This enables users immediately use partitioned results effectively in downstream applications (e.g. RAG architecture apps) without any additional post-processing.languages as an input parameter and marks ocr_languages kwarg for deprecation in pdf, image, and auto partitioning functions. Previously, language information was only being used for Tesseract OCR for image-based documents and was in a Tesseract specific string format, but by refactoring into a list of standard language codes independent of Tesseract, the unstructured library will better support languages for other non-image pipelines and/or support for other OCR engines.UNSTRUCTURED_LANGUAGE env var usage and replaces language with languages as an input parameter to unstructured-partition-text_type functions. The previous parameter/input setup was not user-friendly or scalable to the variety of elements being processed. By refactoring the inputted language information into a list of standard language codes, we can support future applications of the element language such as detection, metadata, and multi-language elements. Now, to skip English specific checks, set the languages parameter to any non-English language(s).xlsx and xls filetype extensions to the skip_infer_table_types default list in partition. By adding these file types to the input parameter these files should not go through table extraction. Users can still specify if they would like to extract tables from these filetypes, but will have to set the skip_infer_table_types to exclude the desired filetype extension. This avoids mis-representing complex spreadsheets where there may be multiple sub-tables and other content.unstructureds NLP internals.unstructured_pytesseract.run_and_get_multiple_output function to reduce the number of calls to tesseract by half when partitioning pdf or image with tesseractunstructured-ingest to write partitioned data from over 20 data sources (so far) to a Delta Table.partition_html would return HTML-elements but now we preserve the format from the input using source_format argument in the partition call.PaddleOCR as an optional alternative to Tesseract for OCR in processing of PDF or Image files, it is installable via the makefile command install-paddleocr. For experimental purposes only.metadata.text_as_html in an element. These changes include:
ENTIRE_PAGE_OCR to specify using paddle or tesseract on entire page OCRcells_to_html doesn't handle cells spanning multiple rows properly (0.5.25)cv2 preprocessing step before OCR step in table transformer (0.5.24)category_depth with default value None.
parent_id on the element's metadata
add_pytesseract_bboxes_to_elements no longer returns nan values. The function logic is now broken into new methods
_get_element_box and convert_multiple_coordinates_to_new_systempartition_image. Problem: partition_pdf allows for passing a model_name parameter. Given the similarity between the image and PDF pipelines, the expected behavior is that partition_image should support the same parameter, but partition_image was unintentionally not passing along its kwargs. This was corrected by adding the kwargs to the downstream call.Update all connectors to use new downstream architecture
Updated documentation: Added back support doc types for partitioning, more Python codes in the API page, RAG definition, and use case.
_detect_filetype_from_octet_stream() function to use libmagic to infer the content type of file when it is not a zip file.clean_ligatures function to expand ligatures in textpartition_html breaks on <br> elements.Removed PIL pin as issue has been resolved upstream
safe_division (0.5.21)Combine entire-page OCR output with layout-detected elements, to ensure full coverage of the page (0.5.19)
xy-cut sorting attemps to sort elements without valid coordinates; now xy cut sorting only works when all elements have valid coordinatesAdds text as an input parameter to partition_xml.
text as an input parameter to partition_xml.partition_xml no longer runs through partition_text, avoiding incorrect splitting
on carriage returns in the XML. Since partition_xml no longer calls partition_text,
min_partition and max_partition are no longer supported in partition_xml.unstructured-inference==0.5.18, change non-default detectron2 classification thresholdelements and bboxes are passed into add_pytesseract_bbox_to_elementsFix test_json to handle only non-extra dependencies file types (plain-text)
test_json to handle only non-extra dependencies file types (plain-text)chunk_by_title to break a document into sections based on the presence of Title
elements.add_pytesseract_bbox_to_elements's (ocr_only strategy) metadata.coordinates.points return type to Tuple for consistency.test_json to handle only non-extra dependencies file types (plain-text)chunk_by_title to break a document into sections based on the presence of Title
elements.extract_image_urls_from_html to extract all img related URL from html text.add_pytesseract_bbox_to_elements's (ocr_only strategy) metadata.coordinates.points return type to Tuple for consistency.Release docker image that installs Python 3.10 rather than 3.8
Remove overly aggressive ListItem chunking for images and PDF's which typically resulted in inchorent elements.
Adds deprecation warning for the file_filename kwarg to partition, partition_via_api, and partition_multiple_via_api.
partition_email and partition_msg to detect if an email is PGP encryped. If
and email is PGP encryped, the functions will return an empy list of elements and
emit a warning about the encrypted content.xy-cut sorting approach in partition_pdf for hi_res and fast strategiespartition_html to respect the order of <pre> tags.partition_pdf_or_image where two partitions were called if strategy == "ocr_only".file_filename kwarg to partition, partition_via_api,
and partition_multiple_via_api.partition raises an error and tells the user to install the appropriate extra if a filetype is detected that is missing dependencies.
partition raises an error and tells the user to install the appropriate extra if a filetype
is detected that is missing dependencies.unstructured-ingest==0.5.15
test_from_image_file in test_layout (0.5.14)entire_page ocr mode for pdfs and imagespartition raises an error and tells the user to install the appropriate extra if a filetype
is detected that is missing dependencies.unstructured-ingest==0.5.15
test_from_image_file in test_layout (0.5.14)entire_page ocr mode for pdfs and imagesAdds ability to reuse connections per process in unstructured-ingest
Bump unstructured-inference==0.5.13:
Bump unstructured-inference==0.5.12:
Update the links and emphasized_texts metadata fields
links and emphasized_texts metadata fieldsinclude_header kwarg to partition_xlsx and change default behavior to Truelinks and emphasized_texts metadata fieldsUpdate partition_csv to always use soupparser_fromstring to parse html text
partition_csv to always use soupparser_fromstring to parse html textpartition_tsv to always use soupparser_fromstring to parse html textmetadata.section to capture epub table of contents dataunique_element_ids kwarg to partition functions. If True, will use a UUID
for element IDs instead of a SHA-256 hash.partition_xlsx to always use soupparser_fromstring to parse html texthtml text parser based on whether the html text contains emojiContent-Distribution: inline and Content-Distribution: attachment with no filename= in the filename itselfpartition_csv to always use soupparser_fromstring to parse html textpartition_tsv to always use soupparser_fromstring to parse html textmetadata.section to capture epub table of contents dataunique_element_ids kwarg to partition functions. If True, will use a UUID
for element IDs instead of a SHA-256 hash.partition_xlsx to always use soupparser_fromstring to parse html texthtml text parser based on whether the html text contains emojiContent-Distribution: inline and Content-Distribution: attachment with no filename= in the filename itselfUpdate table extraction section in API documentation to sync with change in Prod API
element_idAdds --partition-pdf-infer-table-structure to unstructured-ingest.
partition_html to skip headers and footers with the skip_headers_and_footers flag.partition_doc and partition_docx to track emphasized texts in the outputfilter_element_typeshi_respartition_html to track emphasized texts in the outputXMLDocument._read_xml to create <p> tag element for the text enclosed in the <pre> taginclude_tail_text to _construct_text to enable (skip) tail text inclusion_partition_via_api functionpartition_xlsx.file_filename metadata when partitioning file objectEmailAddress for recognizing email address in the textmin_partition logic; makes partitions falling below the min_partition
less likely.Dependencies are now split by document type, creating a slimmer base installation.
Rename "date" field to "last_modified"
Put back useful function split_by_paragraph
split_by_paragraphRemove debug print lines and non-functional code
Add parameter skip_infer_table_types to enable (skip) table extraction for other doc types
skip_infer_table_types to enable (skip) table extraction for other doc typesskip_infer_table_types to enable (skip) table extraction for other doc typesAdditional tests and refactor of JSON detection.
document_to_element_listpartition_html output.convert_to_bytesmin_partition kwarg to that combines elements below a specified threshold and modifies splitting of strings longer than max partition so words are not split.convert_to_bytes--encoding directive to ingestdetect_filetypeimage_metadata property of the PageLayout instance to get the page image info in the document_to_element_listocr_only strategyocr_only strategy.txt, .text, and .tab to list of extensions to check if file
has a text/plain MIME type.partition_doc so it doesn't error with LibreOffice7.requires_dependencies.hi_res as the default strategy value for partition_via_api and partition_multiple_via_apiNLTK now only gets downloaded if necessary.
Fixed auto strategy detected scanned document as having extractable text and using fast strategy, resulting in no output.
auto strategy detected scanned document as having extractable text and using fast strategy, resulting in no output.Allow model used for hi res pdf partition strategy to be chosen when called.
Adjust encoding recognition threshold value in detect_file_encoding
Fix KeyError when isd_to_elements doesn't find a type
Fix _output_filename for local connector, allowing single files to be written correctly to the disk
Fix for cases where an invalid encoding is extracted from an email header.
coordinates attribute of the element's metadata.metadata_filename parameter across all partition functionsconvert_to_datafame grabs all of the metadata fields.detect_file_encodingisd_to_elements doesn't find a type_output_filename for local connector, allowing single files to be written correctly to the diskcoordinates attribute of the element's metadata.Adds include_metadata kwarg to partition_doc, partition_docx, partition_email, partition_epub, partition_json, partition_msg, partition_odt, partition
include_metadata kwarg to partition_doc, partition_docx, partition_email, partition_epub, partition_json, partition_msg, partition_odt, partition_org, partition_pdf, partition_ppt, partition_pptx, partition_rst, and partition_rtfinclude_metadata kwarg to partition_doc, partition_docx, partition_email, partition_epub, partition_json, partition_msg, partition_odt, partition_org, partition_pdf, partition_ppt, partition_pptx, partition_rst, and partition_rtfMore deterministic element ordering when using hi_res PDF parsing strategy (from unstructured-inference bump to 0.5.4)
hi_res PDF parsing strategy (from unstructured-inference bump to 0.5.4)partition_email and partition_msg will now process attachments if process_attachments=True
and a attachment partitioning functions is passed through with attachment_partitioner=partition.Adds a max_partition parameter to partition_text, partition_pdf, partition_email, partition_msg and partition_xml that sets a limit for the size of an
max_partition parameter to partition_text, partition_pdf, partition_email,
partition_msg and partition_xml that sets a limit for the size of an individual
document elements. Defaults to 1500 for everything except partition_xml, which has
a default value of None.hi_res model for pdfs and images is selectable via environment variable.------- as list items.partition_htmlImprovements to string check for leafs in partition_xml.
partition_xml.partition_org for processed Org Mode documents.Adds Google Cloud Service connector
parse_email for partition_eml so that unstructured-api passes the smoke testspartition_email now works if there is no message content"fast" strategy for partition_pdf so that it's able to recursivelyAdds functionality to replace the MIME encodings for eml files with one of the common encodings if a unicode error occurs
MIME encodings for eml files with one of the common encodings if a unicode error occursdetect_file_encodingeml fileshtml_assemble_articles kwarg to partition_html to enable users to capture
control whether content outside of <article> tags is captured when
<article> tags are present.xml attribute on element before looking for pagebreaks in partition_docx.Convert fast startegy to ocr_only for images
.docx and .doc when user or renderer
created page breaks are present.partition_docx to include headers and footers in the output.partition_tsv and associated tests. Make additional changes to detect_filetype.partition_via_api since we now require valid/empty api keysNone instead of 1 when page number is not present in the metadata.
A page number of None indicates that page numbers are not being tracked for the document
or that page numbers do not apply to the element in question..None.Adds functionality to sort elements in partition_pdf for fast strategy
partition_pdf for fast strategy--fast strategy on PDF documentspartition_rst for processed ReStructured Text documents.Allows passing kwargs to request data field for partition_via_api and partition_multiple_via_api
partition_via_api and partition_multiple_via_apidetect_filetype and partition.grpcio import issue on weaviate.schema.validate_schema for python 3.9 and 3.10detectron2 from source in DockerfileUpdate IngestDoc abstractions and add data source metadata in ElementMetadata
strategy parameter down from partition for partition_imagetext/plain MIME typeconvert_office_doc no longers prints file conversion info messages to stdout.partition_via_api reflects the actual filetype for the file processed in the API.Adds an optional encoding kwarg to elements_to_json and elements_from_json
elements_to_json and elements_from_jsonread_txt_file utility function to keep using spooled_to_bytes_io_if_needed for xmlread_txt_file utility function to handle file-like object from URLencoding from partition_pdfNone default for encodingtabulate explicitly to dependenciesmetadata.page_number of pptx filesAdd stage_for_weaviate to stage unstructured outputs for upload to Weaviate, along with a helper function for defining a class to use in Weaviate sche
stage_for_weaviate to stage unstructured outputs for upload to Weaviate, along with
a helper function for defining a class to use in Weaviate schemas.Your coding agent can read these notes before it upgrades. Set up the MCP server →