NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #918 most downloaded on PyPI
A pure-python PDF library capable of splitting, merging, cropping, and transforming PDF files
Last release 4 years ago
no release in 18 months
Ships fairly regularly
a new release about every 2 weeks
Some releases are documented
notes for 32 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
13 years old
66 releases · first in 2013
Nothing published for this version
Deprecate features with PyPDF2==3.0.0
One column per quarter.
Deduplicate extract_text docstring
Consistency changes:
Add support to extract gray scale images
Add AnnotationBuilder.rectangle
Cope with str returned from get_data in cmap
Addition of optional visitor-functions in extract_text()
Add rotation property and transfer_rotate_to_content
Add PageObject.user_unit property
Use pytest.warns() for warnings, and .raises() for exceptions
Fix infinite loop due to Invalid object
Auto-detect RTL for text extraction
Fix errors/warnings on no /Resources within extract_text
Various PdfWriter (Layout, Bookmark deprecation)
BUG: Add PyPDF2.generic to PyPI distribution
BUG: Add PyPDF2.generic to PyPI distribution
Don't check coverage for deprecated code
"with" support for PdfMerger and PdfWriter
Add ability to add hex encoded colors to outline items
Typo in merger deprecation warning message
Add writer.add_annotation, page.annotations, and generic.AnnotationBuilder
Add deprecated EncodedStreamObject functions back until PyPDF2==3.0.0
outline_count property (#1129)Add color and font_format to PdfReader.outlines[i]
build_destination for named destination outlines (#1128)Add support for indexed color spaces / BitsPerComponent for decoding PNGs
Wrong page inserted when PdfMerger.merge is done
Add writer.pdf_header property (getter and setter)
get_form_text_fields does not extract dropdown data
BUG: Forgot to add the internal _codecs subpackage.
BUG: Forgot to add the internal _codecs subpackage.
The highlight of this release is improved support for file encryption (AES-128 and AES-256, R5 only). See #749 for the amazing work of @exiledkingcc 🎊
The highlight of this release is improved support for file encryption (AES-128 and AES-256, R5 only). See #749 for the amazing work of @exiledkingcc 🎊 Thank you 🤗
PdfWriter.get_page: the pageNumber parameter is renamed to page_numberPyPDF2.filters:
PyPDF2.xmp:
PyPDF2.generic:
Apply improvements to _utils suggested by perflint
The 2.2.0 release improves text extraction again via (#969):
The 2.2.0 release improves text extraction again via (#969):
Those changes should mainly improve the text extraction for non-ASCII alphabets, e.g. Russian / Chinese / Japanese / Korean / Arabic.
Mark read_next_end_line as deprecated
PageObject in PyPDF2 root (#960)Mark deprecated code with no-cover
The highlight of the 2.1.0 release is the most massive improvement to the text extraction capabilities of PyPDF2 since 2016 🥳🎊 A very big thank you goes to pubpub-zz who took a lot of time and knowledge about the PDF format to finally get those improvements into PyPDF2. Thank you 🤗💚
In case the new function causes any issues, you can use _extract_text_old
for the old functionality. Please also open a bug ticket in that case.
There were several people who have attempted to bring similar improvements to PyPDF2. All of those were valuable. The main reason why they didn't get merged is the big amount of open PRs / issues. pubpub-zz was the most comprehensive PR which also incorporated the latest changes of PyPDF2 2.0.0.
Thank you to VictorCarlquist for #858 and asabramo for #464 🤗
We introduced a deprecation process that hopefully helps users to avoid unexpected breaking changes.
The 2.0.0 release of PyPDF2 includes three core changes:
We introduced a deprecation process that hopefully helps users to avoid unexpected breaking changes.
overwriteWarnings
parameter. The new behavior is overwriteWarnings=False.ConvertFunctionsToVirtualList was removedformatWarning was removedisInt(obj): Use instance(obj, int) insteadu_(s): Use s directlychr_(c): Use chr(c) insteadbarray(b): Use bytearray(b) insteadisBytes(b): Use instance(b, type(bytes())) insteadxrange_fn: Use range insteadstring_type: Use str insteadisString(s): Use instance(s, str) instead_basestring: Use str insteadPyPDF2.errors:
PyPDF2.pdf (the pdf module) no longer exists. The contents were moved with
the library. You should most likely import directly from PyPDF2 instead.
The RectangleObject is in PyPDF2.generic.Resources, Scripts, and Tests will no longer be part of the distribution
files on PyPI. This should have little to no impact on most people. The
Tests are renamed to tests, the Resources are renamed to resources.
Both are still in the git repository. The Scripts are now in
cpdf. Sample_Code was moved to the docs.For a full list of deprecated functions, please see the changelog of version 1.28.0.
PdfReader.__init__ (#920)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →