NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2528 most downloaded on PyPI
A library that prepares raw documents for downstream ML tasks.
Last release 7 days ago
27 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
3 versions withdrawn
withdrawn after publishing
4 years old
236 releases · first in 2022
One column per quarter.
fix(html): extract definition lists instead of discarding them by @r0h1tb in #4502
<dl>, <dt> and <dd> were mapped to RemovedBlock, so partition_html() dropped every glossary and every Sphinx-generated API reference (each documented function with its parameters and return value) without an error. <dl> is now a list container, each <dd> definition a ListItem and each <dt> term an ordinary text block. The v2 (ontology) parser already kept definition lists.fix(chunking): retain table rowspans across cell splits by @cragwolfe in #4495
rowspan="0" remains scoped to its original table section. Sparse tables with many span expirations are processed without repeatedly scanning every active span.fix: partition multi-section DOCX files in linear time by @linhongyu510 in #4471
Full Changelog: 0.27.6...0.27.8
Doc for text up to 8,192 characters so sentence, word, and part-of-speech tokenization run the spaCy pipeline only once per distinct text.GLOBAL_WORKING_PROCESS_DIR no longer crashes on Windows. Use os.getpid() when the POSIX-only os.getpgid() is unavailable.
Preserve HTML table header semantics. The v1 HTML parser now retains <thead>, <tbody>,
and <tfoot> row groups and preserves <th> cells in Table.metadata.text_as_html. Table text,
nested content extraction, and attribute sanitization are unchanged. Chunking now detects and
repeats eligible v1 HTML header rows by default; repeat_table_headers=False disables header
repetition. Split-table chunk text now treats <br> as a word boundary for all table sources.
Recognize HTML and Markdown loose-list items. A list item containing a single ordinary text block (such as <li><p>text</p></li>) now produces a ListItem, preserving inline annotations and list depth. Multi-paragraph items and specialized blocks retain their existing behavior. Resolves #3499.
partition_doc() and partition_ppt() no longer fail on a document whose name contains multi-byte characters. convert_office_doc() decoded soffice stdout and stderr with a strict UTF-8 decode purely to log them and to check whether stdout was empty. LibreOffice echoes the input path using the console encoding, which on Windows is the locale codepage, so a document whose name or path contains multi-byte characters raised UnicodeDecodeError and aborted a conversion that would otherwise have succeeded. All three decode sites now go through one helper using errors="backslashreplace", which keeps the message pure ASCII -- readable, still loggable by a handler using the locale codepage, and showing the offending bytes. Resolves #3652.
A stray processing instruction no longer crashes HTML partitioning. partition_html (and formats that route through it, such as .md) raised AttributeError: 'lxml.etree._ProcessingInstruction' object has no attribute 'is_phrasing' when the HTML contained a processing-instruction node like a <?xml ...?> declaration. The parser now drops processing instructions at parse time, the same way it already drops comments.
Fix DOCX text_as_html duplicating merged-cell text instead of colspan/rowspan by @qued in #4469
Full Changelog: 0.27.5...0.27.6
text_as_html. A merged cell (gridSpan/vMerge) was repeated into every <td> its merge visually covered, with no colspan/rowspan attribute marking the merge; merged cells are now emitted once, with colspan/rowspan reflecting the true geometry. Since DOCX tables can now carry real spans, table chunking was also made rowspan-aware, so a chunk boundary can no longer split a table in a way that misattributes a spanned cell's rows to the wrong columns.chore(deps): bump claude-code-action to v1 by @william-u10d in #4448
Full Changelog: 0.27.1...0.27.5
auto.partition() re-stamped metadata.filetype on every element with the containing message's type, so content extracted from an attached PDF was labelled message/rfc822. Attachment elements now retain the filetype assigned by the nested partition() call that produced them, including when the containing document is a file-like object with no known file-name.fix: support core metadata 2.5 publishing by @cragwolfe in #4452
Terminology update: Platform -> Pipelines in README by @Paul-Cornell in #4404
Full Changelog: 0.25.0...0.25.2
Speed up HTML element hierarchy reconstruction: elements_to_html() now indexes elements by ID before attaching children, avoiding repeated linear parent scans.
Add lazy chunking entry points: iter_chunk_elements() and iter_chunks_by_title() yield each chunk as it is formed, alongside the list-returning chunk_elements() and chunk_by_title(), which are now defined in terms of them. Same options, same chunks, same order — chunking was already lazy internally and this exposes that pipeline rather than adding a second one. Chunks are no longer accumulated in a list, so a caller that also reads elements lazily holds only the pre-chunk being formed; see the docstrings for two limits on that — iter_chunks_by_title() reads one pre-chunk ahead in order to combine undersized ones, and the default include_orig_elements=True retains source elements (image_base64 payloads included) in every chunk. Options are validated at the call rather than on first advance, and an unknown tokenizer used with max_tokens now raises there too, in both forms.
tokenizer when chunking by max_tokens: "" is not None, so it slipped past the "tokenizer is required" check while still leaving the chunkers without a token counter — the window was then silently measured in characters, making max_tokens=20 mean 20 characters. It now raises the same ValueError as omitting tokenizer altogether.is_json_processable() and is_ndjson_processable() are deprecated : partitioning and file-type detection no longer route through these prefix-sniffing…
partition_json() and partition_ndjson() now handle any valid JSON/NDJSON payload, not just serialized Unstructured output. Arrays (and NDJSON files) of serialized elements keep rehydrating as before; any other valid payload (bare objects, arrays of records, NDJSON lines, scalars) becomes Text elements containing the pretty-printed JSON instead of raising. The schema pre-gates in partition() are removed accordingly, a compact single-line JSON object now detects as FileType.JSON rather than NDJSON (JSON/NDJSON disambiguation examines at most the first 1 MiB of the file), and malformed input still raises ValueError (empty or whitespace-only documents yield no elements). One degraded case: NDJSON whose first record alone exceeds the 1 MiB disambiguation bound now classifies as JSON and fails partition() with ValueError (calling partition_ndjson() directly still handles it). Rehydration is chosen by an explicit shape predicate, with these consequences: an element-shaped payload whose contents cannot be rehydrated (e.g. corrupt metadata) raises ValueError with the underlying error chained, and an array (or NDJSON file) mixing element-shaped and arbitrary items partitions whole as arbitrary JSON - no partial rehydration that silently drops the arbitrary items. An empty JSON object yields one Text containing {} (an empty array yields no elements). One intended routing note: a one-record serialized-element file (a single object, not an array) routed through partition()/detect_filetype() now emits pretty-printed Text with alphabetized keys instead of rehydrating, since rehydration applies only to arrays (direct partition_ndjson() behavior is unchanged).TableChunk elements now rehydrate: elements_from_dicts() (and with it partition_json() and partition_ndjson()) previously dropped serialized TableChunk elements silently because the type is not in the shared element-type map; it is now special-cased like CheckBox. This completes the table-reconstruction feature (#4291), whose reconstruct_table_from_chunks() expects deserialized chunks and now has a deserialization path to feed it. Behavior change: payloads of serialized chunked output containing split tables now return the TableChunk elements (previously omitted from results).is_json_processable() and is_ndjson_processable() are deprecated: partitioning and file-type detection no longer route through these prefix-sniffing helpers. They keep working unchanged for downstream callers - now emitting a DeprecationWarning - and will be removed in a future release.fix: sanitize v2 HTML output to prevent stored XSS ( GHSA-v5mq-3xhg-98m9 ) by @badGarnet in #4394
Full Changelog: 0.24.0...0.24.1
partition_html(html_parser_version="v2"), elements_to_html(), and metadata.text_as_html previously emitted untrusted document markup without output encoding, allowing attacker-controlled content (on* handlers, javascript: links, tag/attribute breakout) to execute when the HTML was viewed. Output is now sanitized — text and attribute values are HTML-escaped, event-handler attributes are dropped, tags/attributes are allowlisted, and URL schemes are filtered (http/https/mailto/tel/relative preserved; data: limited to raster image MIME types on img[src]). Legitimate formatting is unaffected.feat: derive category_depth from heading level in the v2 (ontology) HTML parser by @qued in #4360
Full Changelog: 0.23.1...0.24.0
partition, partition_html, and partition_md now route url= fetches through a single shared helper (unstructured/safe_http.py) instead of ad-hoc requests.get calls. The helper applies an http/https scheme allowlist, a hostname denylist with IDNA normalization, address validation performed at connect time, manual redirect handling with per-hop re-validation (dropping credential material on cross-origin hops), refusal of proxied requests, and a default (connect, read) timeout. Behavior change: fetches that resolve to non-routable, loopback, or link-local addresses are now rejected by default. Set UNSTRUCTURED_ALLOW_PRIVATE_URL=1 (or pass allow_private=True) to opt out for controlled local usage.feat: extract filled AcroForm field text in PDF partitioning by @badGarnet in #4372
Full Changelog: 0.23.0...0.23.1
fast and hi_res strategies and emitted as elements alongside the content-stream text.any(extracted_to_keep), which evaluated False when the only kept extracted region was at index 0. On single-region pages (e.g. a PDF whose only text is one filled form field) this left a duplicate element; the guard now checks the array size.fix: stop decimating embedded text on dense PDF pages by @badGarnet in #4368
Full Changelog: 0.22.32...0.23.0
enrichment_origins metadata field for per-attribute model provenance: ElementMetadata gains a serialized enrichment_origins field mapping a written attribute name (e.g. text, text_as_html, embeddings) to a list of records {"type", "provider", "model"}, in application order. Enrichment producers stamp which model wrote (or contributed to) each attribute; authoring enrichments overwrite the list while additive ones append, preserving the prior author. A new ConsolidationStrategy.DICT_LIST_UNIQUE merges these dicts across elements during chunking (union keys, concatenate then dedupe records, preserving first-seen order).fix(hi_res): recover text inside PDF figure overlays by @qued in #4363
Full Changelog: 0.22.31...0.22.32
get_text (e.g. LTTextBox), and extract_text_objects only collected LTTextLine. Text held as loose LTChars inside an LTFigure - for example text drawn into a figure/XObject overlay rather than the main content stream - was dropped from the output. hi_res now groups such loose characters into text lines, inserting spaces on wide inter-character gaps and skipping hidden (render mode 3) and rotated characters.fix: rename isolate_tables chunking option to isolate_table by @badGarnet in #4355
isolate_tables chunking option to isolate_table by @badGarnet in #4355Full Changelog: 0.22.30...0.22.31
isolate_tables chunking option to isolate_table: the option added in 0.22.30 has been renamed for naming consistency. Callers passing isolate_tables= must update to isolate_table=.feat: add option for table chunking by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4354
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.29...0.22.30
isolate_tables to basic/title chunking options. Defaults to True (the post-#4307 behavior: Table/TableChunk elements always staged alone). Set to False to allow tables to share pre-chunks with adjacent non-table elements and be combined by PreChunkCombiner.fix: handle text too long for spacy issue by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4353
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.28...0.22.29
spacy limit: add a guard against calling spacy tokenizer with very long text. Now long texts are truncated to fit under the character limit.fix: chunking dropping table content by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4352
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.27...0.22.28
HtmlTable compactification previously cleared every element's .tail, silently dropping real text between inline children. Pure-whitespace tails are still removed, but tails carrying content are now kept with internal whitespace collapsed.fix: ndjson file type detection by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4349
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.26...0.22.27
is_ndjson_processable previously returned True for any text starting with {, so .json and .ipynb files containing a single multi-line JSON object (e.g. Jupyter notebooks) were routed to partition_ndjson, which then crashed in its splitlines()-based parser.Reject oversized PDF renders before bitmap allocation by @CyMule in https://github.com/Unstructured-IO/unstructured/pull/4345
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.23...0.22.26
table_extraction_method field to ElementMetadata to track which algorithm produced a table (grid, tatr, vlm). Propagated from LayoutElement during PDF partitioning.fix: first table chunk preserve col/row span by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4343
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.22...0.22.23
colspan/rowspan in first table chunk headers: HtmlTable compactification no longer strips colspan and rowspan attributes from table cells. Previously, the first TableChunk lost merged-cell structural information while continuation chunks retained it (via the source-HTML path used for repeated headers), yielding inconsistent header layout across a split table.Replace PyPI opencv wheels with ffmpeg-free builds in Docker image: After uv sync, the Dockerfile now substitutes all PyPI opencv-python variants with
uv sync, the Dockerfile now substitutes all PyPI opencv-python variants with a source-built opencv-contrib-python-headless wheel compiled with WITH_FFMPEG=OFF, eliminating 14 bundled ffmpeg CVEs. The contrib-headless variant is a strict superset of the cv2 API (core + contrib modules, no GUI) so a single wheel replaces opencv-python, opencv-python-headless, and opencv-contrib-python.feat: add option to skip table chunking by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4338
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.20...0.22.21
skip_table_chunking to basic/title chunking options. When True, Table elements are passed through unchanged without being split into TableChunk elements, regardless of their size. Defaults to False to preserve existing behavior.fix(deps): upgrade vulnerable transitive dependencies [security] by @utic-github-cicd-token-generator[bot] in https://github.com/Unstructured-IO/unstr…
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.18...0.22.20
detect_vertical field to PDFMinerConfig and auto-enable it when rendered pages have /Rotate metadata, so pdfminer groups rotated text into proper words instead of per-character regionsfix(chunking): preserve semantic headers in carried table chunks by @cragwolfe in https://github.com/Unstructured-IO/unstructured/pull/4313
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.16...0.22.18
ingest-test-fixtures-update-pr CI job also update the markdown versions of the fixtures.data-page-number attributes from ancestor elements and includes the page number in element metadata, consistent with the v2 parser behavior.security: fix(deps): upgrade vulnerable transitive dependencies [security]
element_to_md / elements_to_md): New keyword-only formula_markdown_style ("auto", "display_math", "plain"; default "auto"). In "auto", display math ($$ ... $$) is used only when the text looks like notation (heuristic score) and contains no $/$$ (avoids breaking Markdown and noisy OCR captions). "display_math" wraps whenever safe (still falls back to plain if $ would corrupt fences). "plain" emits text only. Optional normalize_formula (default True) maps common Unicode operators to LaTeX-like tokens; normalize_formula stays before keyword-only options so positional encoding / no_group_by_page callers are unchanged. Unicode √ is never mapped to \\sqrt{}. Module constants: FORMULA_MARKDOWN_AUTO, FORMULA_MARKDOWN_DISPLAY_MATH, FORMULA_MARKDOWN_PLAIN._render_pdf_pages and delegate to unstructured-inference's convert_pdf_to_image (which already has lazy per-page rendering). Peak memory for path_only=True drops from O(n_pages) to O(1 page) — 97% reduction on a 100-page PDF. Bumps inference dep to >=1.6.2.standardize_quotes: Replace loop-based character replacement with a single str.translate() call using a pre-computed translation table. Also fixes a pre-existing bug where left smart quotes were never normalized due to duplicate dictionary keys.mem: exclude unused spaCy pipeline components to reduce model memory by @KRRT7 in https://github.com/Unstructured-IO/unstructured/pull/4296
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.10...0.22.12
fix(chunking): preserve nested table structure in reconstruction by @cragwolfe in https://github.com/Unstructured-IO/unstructured/pull/4301
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.6...0.22.10
repeat_table_headers to basic/title chunking options and table chunking internals so leading header rows are detected once and carried forward when large tables spill across multiple chunks.fix(deps): Update security updates [SECURITY] by @utic-renovate[bot] in https://github.com/Unstructured-IO/unstructured/pull/4303
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.22.4...0.22.6
feat: custom fallback for language detection by @claytonlin1110 in https://github.com/Unstructured-IO/unstructured/pull/4238
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.21.2...0.21.5
Self-install pinned spaCy model at runtime with SHA256 verification: Replace the en-core-web-sm direct URL dependency in pyproject.toml with the insta
en-core-web-sm direct URL dependency in pyproject.toml with the installer library. The spaCy model is now downloaded and installed on first use with hash verification, removing the need for [tool.uv.sources] and making the install more portable.bump version by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4257
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.21.0...0.21.1
Replace NLTK with spaCy to remediate CVE-2025-14009: NLTK's downloader uses zipfile.extractall() without path validation, enabling RCE via malicious p…
zipfile.extractall() without path validation, enabling RCE via malicious packages (CVSS 10.0, no patch available). spaCy models install as pip packages, eliminating the vulnerable downloader entirely.fix: set max decompressed size for elements JSON by @qued in https://github.com/Unstructured-IO/unstructured/pull/4244
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.20.6...0.20.8
wrapt so it is compatible with opentelemetry-instrumentation-httpxAutomate pypi publishing by @PastelStorm in https://github.com/Unstructured-IO/unstructured/pull/4239
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.20.1...0.20.6
Add automated PyPI publishing: new release.yml GitHub Actions workflow triggers on GitHub release, builds the package with uv build, publishes to PyPI
release.yml GitHub Actions workflow triggers on GitHub release, builds the package with uv build, publishes to PyPI via pypa/gh-action-pypi-publish, and uploads to Azure Artifacts via twineuv sync --frozen with uv sync --locked across all CI workflows, Dockerfile, and Makefile to fail fast on stale lockfiles--no-sync to all uv run and uv build commands that follow a prior uv sync step to prevent implicit re-syncingfeat: put pdfium call behind a threadlock by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4211
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.18.31...0.18.32
Feat: patch pdfminer and use rendermode to detect invisible text by @badGarnet in https://github.com/Unstructured-IO/unstructured/pull/4158
_get_optimal_value_for_bbox by 2,883% by @aseembits93 in https://github.com/Unstructured-IO/unstructured/pull/4181_DocxPartitioner._style_based_element_type by 593% by @aseembits93 in https://github.com/Unstructured-IO/unstructured/pull/4179Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.18.28...0.18.31
max_tokens, new_after_n_tokens, and tokenizer parameters to chunk_by_title() and chunk_elements() for chunking by token count instead of character count. Uses tiktoken for token counting. Install with pip install "unstructured[chunking-tokens]". (fixes #4127)ALLOW_PANDOC_NO_SANDBOX=true env var is set (fixes #3997)coordinates=True causing TypeError in hi_res PDF processing: Filter out coordinates and coordinate_system from kwargs before passing to add_element_metadata() to prevent conflict with explicit parameters (fixes #4126)<pre> elements now generate CodeSnippet elements instead of Text, and chunking preserves internal whitespace for code snippets. (fixes #4095)Comment no-ops in zoom_image (codeflash)
zoom_image (codeflash)sentence_count (codeflash)_PartitionerLoader._load_partitioner (codeflash)detect_languages (codeflash)contains_verb (codeflash)get_bbox_thickness (codeflash)Pin deltalake<1.3.0 to fix ARM64 Docker builds (1.3.0 missing Linux ARM64 wheels)
deltalake<1.3.0 to fix ARM64 Docker builds (1.3.0 missing Linux ARM64 wheels)Security update: Bumped dependencies to address security vulnerabilities
OCRAgentTesseract.extract_word_from_hocr (codeflash)Update save_elements unit test to check crop box padding behavior
unstructured-inference to 1.1.2 to address CVEsImprove the VoyageAI integration
Prevent path traversal in email MSG attachment filenames Fixed a security vulnerability (GHSA-gm8q-m8mv-jj5m) where malicious attachment filenames con…
partition_msg functionsClarifai dependency as it is no longer usedSetup Codeflash Github Actions to optimize all future code by @misrasaurabh1 in https://github.com/Unstructured-IO/unstructured/pull/4082
group_broken_paragraphs by 30% by @aseembits93 in https://github.com/Unstructured-IO/unstructured/pull/4088ElementHtml._get_children_html by 234% by @aseembits93 in https://github.com/Unstructured-IO/unstructured/pull/4087Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.18.14...0.18.15
Python 3.12/3.13: CVE-2025-8194, GHSA-v594-44hm-2j7p
Speed up function sentence_count by 59% (codeflash)
Speed up function check_for_nltk_package by 111% (codeflash)
Speed up function under_non_alpha_ratio by 76% (codeflash)
Parse a wider variety of date formats in email headers The partition_email function is now more robust to non-standard date formats, including ISO-860
Parse a wider variety of date formats in email headers The partition_email function is now more robust to non-standard date formats, including ISO-8601 dates with "Z" suffixes. This prevents ValueError exceptions when partitioning emails with these date formats.
add '|' as a delimiter in csv files by @jiajun-unstructured in https://github.com/Unstructured-IO/unstructured/pull/4059
type + add coverage by @MaksOpp in https://github.com/Unstructured-IO/unstructured/pull/4068Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.18.10...0.18.11
charset-normalizer library for encoding detection Previously we had both chardet and charset-normalizer as dependencies. We are dropping chardet and only using charset-normalizer.<input> mapping in HTML transformations Bare <input> elements are now classified by their type attribute (checkbox → Checkbox, radio → RadioButton, others → FormFieldValue).Convert elements to markdown for output Added function to convert elements to markdown format for easy viewing.
Add language detection for PDFs Add document and element level language detection to PDFs.
text_as_html for Table element now keeps both input and img tag's class attribute Previously in partition HTML any tag inside a table is stripped of its class attribute. Now this attribute is preserved for both input and img tag in the table element's metadata.text_as_html.Improved epub partition errors EPUB partition will now produce new type of error on unprocessable files.
TableChunk for the string value of the field type when serializing elements of type TableChunk, rather than using the value Table.Bump dependencies and remove lingering Python 3.9 artifacts Cleaned up some references to 3.9 that were left When we dropped Python 3.9 support.
text_as_html for Table element now keeps img tag's class attribute Previously in partition HTML any tag inside a table is stripped of its class attribute. Now this attribute is preserved for img tag in the table element's metadata.text_as_html.chore: bump pillow to address a CVE by @awalker4 in https://github.com/Unstructured-IO/unstructured/pull/4045
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.18.2...0.18.3
fix [NEX-49] : Fix TypeError for empty HTML content by @yuming-long in https://github.com/Unstructured-IO/unstructured/pull/4032
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.18.1...0.18.2
tc_at_grid_offset and raised ValueError: no tc element at grid_offset=X.partition_md reads the file as utf-8 previously. Now it uses read_txt_file that reads file with detected encoding.UncategorizedText or as the nested structure like Title. Now they are properly partitioned as Header and Footer element types.Add DocumentData element type This is helpful in scenarios where there is large data that does not make sense to represent across each element in the
encoding property of the _CsvPartitioningContext is now properly used.Add image_url of images in html partitioner tags with non-data content include a new image_url metadata field with the content of the src attribute.
Add image_url of images in html partitioner <img> tags with non-data content include a new image_url metadata field with the content of the src attribute.
Use lxml instead of bs4 to parse hOCR data. lxml is much faster than bs4 given the hOCR data format is regular (garanteed because it is programatically generated)
bump numpy to >2. And upgrade paddlepaddle, unstructured-paddleocr, onnx so they are compatible with numpy>2.
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.17.0...0.17.2
feat: include images when partitioning html by @ryannikolaidis in https://github.com/Unstructured-IO/unstructured/pull/3945
Full Changelog: https://github.com/Unstructured-IO/unstructured/compare/0.16.25...0.17.0
Add support for images in html partitioner <img> tags will now be parsed as Image elements. When extract_image_block_types includes Image and extract_image_block_to_payload=True then the image_base64 will be included for images that specify the base64 data (rather than url) as the source.
Use kwargs instead of env to specify ocr_agent and table_ocr_agent for hi_res strategy.
stop using PageLayout.elements to save memory and cpu cost. Now only use PageLayout.elements_array throughout the partition, except when analysis=True where the drawing logic still uses elements.
Fixes filetype detection for jsons passed as byte streams - Now it prioritizes magic mimetype prediction over file extension when detecting filetypes
Support dynamic partitioner file type registration. Use create_file_type to create new file type that can be handled in unstructured and register_part
Support dynamic partitioner file type registration. Use create_file_type to create new file type that can be handled
in unstructured and register_partitioner to enable registering your own partitioner for any file type.
extract_image_block_types now also works for CamelCase elemenet type names. Previously NarrativeText and similar CamelCase element types can't be extracted using the mentioned parameter in partition. Now figures for those elements can be extracted like Image and Table elements
use block matrix to reduce peak memory usage for pdf/image partition.
Fixes detect_filetype when SpooledTemporaryFile is passed. Previously some random name would get assigned to the file and the function raised error.
Fix open CVES in and bump dependencies
Your coding agent can read these notes before it upgrades. Set up the MCP server →