NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4237 most downloaded on PyPI
A package to use AWS Textract services.
Last release 1 months ago
11 Aug 2026
Ships unpredictably
gaps range from 2 weeks to 1.3 years
Nearly every release is documented
notes for 50 of 52 stable releases
2 versions withdrawn
withdrawn after publishing
4 years old
54 releases · first in 2022
One column per quarter.
Replace editdistance with rapidfuzz for Levenshtein distance by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/452
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.9.2...v1.10.0
Ignore KV elements in LAYOUT_LIST by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/425
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.9.1...v1.9.2
Fix s3 client instantiation in _get_document_images_from_path by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/423
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.9.0...v1.9.1
Add function to split larger element at table insertion by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/417
export_kv_to_pandas by @Chuukwudi in https://github.com/aws-samples/amazon-textract-textractor/pull/405Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.8.5...v1.9.0
Fix bug in convert that caused an exception on empty pages.
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.8.4...v1.8.5
Add check for None bounding boxes for AnalyzeExpense by @Belval
Document.export_kv_to_csv() by @ChuukwudiFull Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.8.3...v1.8.4
Id in html output by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/386
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.8.2...v1.8.3
Fix pypdfium2 failing to parse PDFs in bytearray format by @Belval
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.8.1...v1.8.2
Fix .to_markdown() raising an exception on missing local config by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/381
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.8.0...v1.8.1
Add HTML table linearization format that uses merged cells information for colspan and rowspan
colspan and rowspanLAYOUT_FOOTER and LAYOUT_ENTITY<html><body>...</body></html> to the output when calling Document.to_html()pypdfium2 for PDF rasterization when available instead of pdf2image. This allows for better portability as the former does not have a dependency on OS libraries and should work out of the box with Lambda and SageMaker.s3_output_path from the synchronous functions as s3_output_path is not a supported parameter for the Textract Synchronous APItextractor.py functions which will no longer raise RegionMismatchError (which is however kept in textractor.exceptions for backward compatibility.confidence_score from KeyValue entities in favour of _confidence which is used for all other entities.Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.12...v1.8.0
Fix issue where tables linearized to plaintext that contained merge cells would duplicate the text over the entire table.
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.11...v1.7.12
Add figure layout prefix and suffix by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/362
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.10...v1.7.11
Use AWS_REGION and AWS_DEFAULT_REGION environment variables in Textractor when available
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.9...v1.7.10
Set JPEG compression parameters by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/342
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.8...v1.7.9
Handle None Relationships when parsing LAYOUT_FIGURE
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.7...v1.7.8
Handle None bounding box when parsing Queries by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/340
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.6...v1.7.7
Add CITATION.cff by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/332
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.5...v1.7.6
Make KeyValue.key an EntityList by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/320
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.4...v1.7.5
Fix table title .get_text() by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/314
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.3...v1.7.4
Table linearization improvements by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/313
Table linearization improvements by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/313
.get_text(), .to_html() and .to_markdown() functions to Linearizable which is now implemented by Document, Page, DocumentEntity and EntityListHTMLLinearizationConfig and MarkdownLinearizationConfig as pre-configured TextLinearizationConfigTextLinearizationConfig
duplicate_text_in_merged_cells duplicates the text in merge cells to preserve row-level alignmenttable_flatten_headers combines multi-row headers into a single row, duplicating the merged cells horizontally as neededtable_tabulate_remove_extra_hyphens removes extra hyphens '-' in markdown tables to reduce context lengthmax_number_of_consecutive_spaces defines the maximum number of contiguous whitespace characters, similar to max_number_of_consecutive_new_linesFixes:
table_column_separator being hardcoded as '\t'table_row_separator being hardcoded as '\n'Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.2...v1.7.3
Fix for page objects not always having an image attached, causing an exception on .visualize()
.visualize()Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.1...v1.7.2
Fix issue where a table within a container layout could be duplicated in the .get_text() output.
.get_text() output.Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.7.0...v1.7.1
Loosen XlsxWriter version constraints by @mdscruggs in https://github.com/aws-samples/amazon-textract-textractor/pull/292
textractor.__version__ to allow easier identification of the installed Textractor version in codelinearize_table and linearize_key_value from TextLinearizationConfig as both were not useds3_output_path parameter from analyze_expense as the API does not support outputting to S3Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.6.1...v1.7.0
## What's new - Fix bug in table to markdown Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.6.0...v1.6.1
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.6.0...v1.6.1
Fix selection elements in table by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/289
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.5.0...v1.6.0
Nothing published for this version
Add GetResult from S3 in LazyDocument
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.4.5...v1.5.0
Fix missing words in get_text_and_words by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/270
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.4.4...v1.4.5
Add page_layout property to Page object by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/268
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.4.3...v1.4.4
Raise exception on export-to-markdown without pandas by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/261
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v.1.4.2...v1.4.3
Fix: #257, adjust versions for caller by @schadem in https://github.com/aws-samples/amazon-textract-textractor/pull/258
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.4.1...v.1.4.2
Fix signature token not being added to the linearized text
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.4.0...v1.4.1
Version 1.4.0 introduces text linearization, powered by the new Amazon Textract Layout feature. Layout prediction enables the conversion of documents
Version 1.4.0 introduces text linearization, powered by the new Amazon Textract Layout feature. Layout prediction enables the conversion of documents to text while preserving their reading order. See our layout prediction notebook for an example.
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.7...v1.4.0
Fix STRUCTURED and SEMI_STRUCTURED table types not being properly parsed
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.6...v1.3.7
Fix Levenshtein distance not returning the most similar key by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/245
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.5...v1.3.6
## What's Changed * Add support for Python 3.11 Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.4...v1.3.5
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.4...v1.3.5
Pass region_name as argument to Textract client when provided by @alanmohan in https://github.com/aws-samples/amazon-textract-textractor/pull/229
region_name as argument to Textract client when provided by @alanmohan in https://github.com/aws-samples/amazon-textract-textractor/pull/229Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v.1.3.3...v1.3.4
Fix bbox denormalization breaking visualizations by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/242
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.2...v.1.3.3
Fix assertion error on empty line item rows by @vinyasmusic in https://github.com/aws-samples/amazon-textract-textractor/pull/224
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.1...v1.3.2
fix: ordering of paginated files on S3 OutputConfig by @schadem in https://github.com/aws-samples/amazon-textract-textractor/pull/208
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.3.0...v1.3.1
Significant improvement to response validation time by @krzim in https://github.com/aws-samples/amazon-textract-textractor/pull/203
amazon-textract-response-parser has already dropped supportFull Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.2.0...v1.3.0
Add support for titles, footers and cell types by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/197
Update examples.rst by @ThomasDelteil in https://github.com/aws-samples/amazon-textract-textractor/pull/183
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.1.0...v1.1.1
Add new lambda layer with pandas by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/177
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.24...v1.1.0
Add lambda and signature detection example by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/157
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.23...v1.0.24
Fix XlsxWriter raising an exception on overlapping merged cells by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/153
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.22...v1.0.23
Use editdistance instead of pyxDamerauLevenshtein by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/144
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.21...v1.0.22
Add support for Signatures in the CLI
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.20...v1.0.21
Fix word ordering by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/137
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.18...v1.0.20
Fix start_analyze_expense raising exception by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/124
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.17...v1.0.18
Fix wrong parameter name in Session instantiation by @kkourmousis in https://github.com/aws-samples/amazon-textract-textractor/pull/115
Full Changelog: https://github.com/aws-samples/amazon-textract-textractor/compare/v1.0.16...v1.0.17
Textractor refactoring by @Belval in https://github.com/aws-samples/amazon-textract-textractor/pull/108
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →