NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #404 most downloaded on PyPI
PDF parser and analyzer
Last release 9 months ago
07 Jan 2026
Release timing varies
gaps range from 1 weeks to 1.1 years
Most releases are documented
notes for 29 of 37 stable releases
Nothing withdrawn
no release was ever pulled
12 years old
37 releases · first in 2014
Support cmap types 6, 10 and 12
Eliminated arbitrary code execution vulnerability ( CVE-2025-64512 ) by replacing pickle CMap storage with json - users with custom pickle CMaps can u…
One column per quarter.
tools/convert_cmaps_to_json.py to convert to JSON format (#1172))Support for colored and uncolored tiling patterns per ISO 32000
Support for colored and uncolored tiling patterns per ISO 32000
Refuse to execute circular references to content streams (including Form XObjects)
Support for extracting images with TIFF predictor
Support for extracting images with TIFF predictor
TypeError when parsing font width with indirect object references
TypeError when parsing font width with indirect object references (#1098)ValueError when loading xref with invalid position or generation numbers that cannot be parsed as int (#1099)TypeError when parsing font bbox with incorrect values (#1103)ValueError on incorrect stream lengths for ASCII85 data (#1112)Reduce memory overhead on runlength encoding by using lists
pyproject.toml instead of setup.py (#1028)TypeError when CID character widths are not parseable as floats (#1001)TypeError raised by extract_text method with compressed PDF file (#1029)PSBaseParser can't handle tokens split across end of buffer (#1030)TypeError when CropBox is an indirect object reference (#1004)bytes is bytes where it counts (#1069)Deprecated tools, functions and classes
PDFObjRef (#972)TypeError when corrupt PDF object reference cannot be parsed as int (#972)])TypeError when corrupt PDF literal cannot be converted to str (#978)ValueError when corrupt PDF specifies a negative xref location (#980)ValueError when corrupt PDF specifies an invalid mediabox (#987)RecursionError when corrupt PDF specifies a recursive /Pages object (#998)TypeError when corrupt PDF specifies text-positioning operators with invalid values (#1000)TypeError when parsing object reference as mediabox (#1082)Fuzzing harnesses for integration into Google's OSS-Fuzz (949)
END_KEYWORD before end-of-stream are parsed (#885)ValueError wrong error message when specifying codec for text output (#902)apply_png_predictor by using lists (#912)ValueError when extracting images, due to breaking changes in Pillow
flake8 failures (#921)ValueError when bmp images with 1 bit channel are decoded (#773)ValueError when trying to decrypt empty metadata values (#766)TypeError when getting default width of font (#720)TypeError in cmapdb.py when parsing null characters (#768)ValueError when extracting images, due to breaking changes in Pillow (#827)if __name__ == "__main__" where it was only intended for testing purposes (#756)ValueError when extracting images, due to breaking changes in Pillow
ValueError when bmp images with 1 bit channel are decoded (#773)ValueError when trying to decrypt empty metadata values (#766)TypeError when getting default width of font (#720)TypeError in cmapdb.py when parsing null characters (#768)ValueError when extracting images, due to breaking changes in Pillow (#827)if __name__ == "__main__" where it was only intended for testing purposes (#756)Ignoring (invalid) path constructors that do not begin with m
IndexError when handling invalid bfrange code map in CMap
IndexError when handling invalid bfrange code map in
CMap (#731)TypeError in lzw.py when self.table is not set (#732)TypeError in encodingdb.py when name of unicode is not
str (#733)TypeError in HTMLConverter when using a bytes fontname (#734)Export type annotations from pypi package per PEP561
LTLayoutContainer.group_textboxes that returned some text lines out of order (#659)pdf2txt.py --boxes-flow=disabled (#682)PDFNoValidXRef is raised and fallback is True (#684)Add support for PDF 2.0 (ISO 32000-2) AES-256 encryption
KeyError when 'Encrypt' but not 'ID' present in trailer (#594)PermissionError when creating temporary filepaths on windows when running tests (#484)AttributeError when dumping a TOC with bytes destinations (#600).paint_path logic for handling single line segments and extracting point-on-curve positions of Beziér path commands (#530)UnboundLocalError when a bad --output-type is used (#610)TypeError when using TagExtractor with non-string or non-bytes tag values (#610)io.TextIOBase as the file to write to (#616)Option to disable boxes flow layout analysis when using pdf2txt
pathlib.PurePath in open_filename (#491)high_level functions (#475).paint_path logic for handling non-rect quadrilaterals and decomposing complex paths (#473)Rename PDFTextExtractionNotAllowedError to PDFTextExtractionNotAllowed to revert breaking change
Support for painting multiple rectangles at once
Python3 shebang line to script in tools (408
Allow boxes_flow LAParam to be passed as None, validate the input, and update documentation
extract_text and extract_pages (#392)word_margin and char_margin (#407)Nothing published for this version
Removed samples/issue-00152-embedded-pdf.pdf because it contains a possible security thread; a javascript enabled object
Interpret two's complement integer as unsigned integer
Enforce pep8 coding style by adding flake8 to CI
Wrong order of text box grouping introduced by PR #315
The argument _py2_no_more_posargs because Python2 is removed on January , 2020 (#328 and #307)
LTLayoutContainer.group_textboxes for a significant speed up in layout analysis (#315)Support for Python 2 is dropped at January 1st, 2020
six.iteritems() instead of dict().iteritems() to ensure Python2 and Python3 compatibility (#274)Use argparse instead of replace deprecated getopt
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →