NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #485 most downloaded on PyPI
Plumb a PDF for detailed information about each char, rectangle, and line.
Last release 3 months ago
15 Jun 2026
Ships fairly regularly
a new release about every 3 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
11 years old
76 releases · first in 2015
Upgrade pdfminer.six from 20251230 to 20260107 .
One column per quarter.
Add edge_min_length_prefilter table setting for initial edge filtering. Lowering this setting enables capturing small edge segments (e.g., dashed line
edge_min_length_prefilter table setting for initial edge filtering. Lowering this setting enables capturing small edge segments (e.g., dashed lines) that would be filtered out with the default minimum length of 1. Raising this setting would be less common but plausible. (h/t @bronislav). (#1274).pdfminer.six from 20250506 to 20251107 (h/t @henry-renner-v). (0079187)Add access to Page.trimbox , Page.bleedbox , and Page.artbox (h/t @samuelbradshaw ). ( #1313 + 7e364e6 )
Page.trimbox, Page.bleedbox, and Page.artbox (h/t @samuelbradshaw). (#1313 + 7e364e6)pdfminer.six from 20250327 to 20250506. (4c7e092)stroking_pattern and non_stroking_pattern object attributes, due to changes in pdfminer.six. (4c7e092)Upgrade pdfminer.six from 20231228 to 20250327 ( 3fcb493 + 12a73a2 )
pdfminer.six from 20231228 to 20250327 (3fcb493 + 12a73a2)use_text_flow=True text extraction (h/t @samuelbradshaw) (#1279 + e15ed98)Add --format text options to CLI (in addition to previously-available csv and json ) (h/t @brandonrobertz ).
--format text options to CLI (in addition to previously-available csv and json) (h/t @brandonrobertz). (#1235)raise_unicode_errors: bool parameter to pdfplumber.open() to allow bypassing UnicodeDecodeErrors in annotation-parsing and generate warnings instead (h/t @stolarczyk). (#1195)name property to image objects (h/t @djr2015). (#1201)PageImage.debug_tablefinder(...) so that its main keyword argument is named the same (table_settings=) as other related Page methods (h/t @stolarczyk). (#1237)Fix one type hint so that it doesn't throw error on Python 3.8 (h/t @andrekeller ).
Add Table.columns , analogous to Table.rows (h/t @Pk13055 ). ( #1050 + d39302f )
Table.columns, analogous to Table.rows (h/t @Pk13055). (#1050 + d39302f)Page.extract_words(return_chars=True), mirroring Page.search(..., return_chars=True); if this argument is passed, each word dictionary will include an additional key-value pair: "chars": [char_object, ...] (h/t @cmdlineluser). (#1173 + 1496cbd)pdfplumber.open(unicode_norm="NFC"/"NFD"/"NFKC"/NFKD"), where the values are the four options for Unicode normalization (h/t @petermr + @agusluques). (#905 + 03a477f)pdfplumber.repair(...) passes to Ghostscript's -dPDFSETTINGS parameter, from prepress to default, and make that setting modifiable via .repair(setting=...), where the value is one of "default", "prepress", "printer", or "ebook" (h/t @Laubeee). (#874 + 48cab3f)mediabox does not begin at (0,0) (h/t @wodny). (#1181 + 9025c3f + 046bd87).annots/.hyperlinks from CroppedPage (due to missing .rotation and .initial_doctop attributes) (h/t @Safrone). (#1171 + e5737d2)Page.crop(...) was not cropping .annots/.hyperlinks (h/t @Safrone). (#1171 + 22494e8).annots on CroppedPages. (0bbb340 + b16acc3)Page.get_attr(...) so that it fully resolves references before determining whether the attribute's value is None (h/t @zzhangyun + @mkl-public). (#1176 + c20cd3b)Add extra_attrs parameter to .dedupe_chars(...) to adjust the properties used when deduplicating (h/t @QuentinAndre11 ).
extra_attrs parameter to .dedupe_chars(...) to adjust the properties used when deduplicating (h/t @QuentinAndre11). (#1114)flake8, pytest, and pytest-cov — and add setuptools and py as explicit dev requirements (for Python 3.12).Fix .open(..., repair=True) subprocess args (to avoid stderr being captured)
Deprecate vertical_ttb, horizontal_ltr in favor of char_dir and char_dir_rotated.
Summary: More control over the {left-to-right, right-to-left, top-to-bottom, bottom-to-top} direction that pdfplumber reads/writes text (many thanks to @afriedman412 for the idea and prototype in #1040), plus upgrading to pdfminer.six's latest release (which provides more detailed paths for curves), and some fixes.
{line,char}_dir{,rotated,render} params, to provide better support for non–top-to-bottom, left-to-right text (h/t @afriedman412). (850fd45)curve["path"] and curve["dash"], thanks to pdfminer.six upgrade (see below). (1820247)pdfminer.six from 20221105 to 20231228. (cd2f768)word["direction"] from {1,-1} to {"ltr","rtl","ttb","btt"}. (850fd45)vertical_ttb, horizontal_ltr in favor of char_dir and char_dir_rotated.(850fd45)Add x_tolerance_ratio parameter to extract_text and similar functions, to account for text size when spacing characters (instead of a fixed number of
x_tolerance_ratio parameter to extract_text and similar functions, to account for text size when spacing characters (instead of a fixed number of pixels) (h/t @afriedman412). (#1041)Page.structure_tree (h/t @dhdaines). (#963)repair.py (h/t @echedey-ls). (#1032)Page.close() method, have PDF.close() close all pages as well, and improve relevant documentation (h/t @luketudge). (#1042)force_mediabox parameter to Page.to_image(...). (#1054)Page.get_textmap caching to allow for extra_attrs=[...], by preconverting list kwargs to tuples. (#1030)pypdfium2.PdfDocument in get_page_image (h/t @dhdaines). (#1090)PDFPageAggregatorWithMarkedContent.tag_cur_item, check self.cur_item._objs length before trying to access [-1]. (4f39d03)Add support for marked-content sequences, represented by mcid and tag attributes on char/rect/line/curve/image objects (h/t @dhdaines).
mcid and tag attributes on char/rect/line/curve/image objects (h/t @dhdaines). (#961)gs_path argument to pdfplumber.open(...) and pdfplumber.repair(...), to allow passing a custom Ghostscript path to be used for repairing. (#953)use_text_flow in extract_text (h/t @dhdaines). (#983)Add PDF.path: A Path object for PDFs loaded by passing a path (unless repair=True), and None otherwise. (30a52cb + #948)
Add antialias boolean parameter to Page.to_image(...) and associated methods (h/t @cmdlineluser).
A simple release:
antialias boolean parameter to Page.to_image(...) and associated methods (h/t @cmdlineluser). (7e28931)Normalize color representation to tuple[float|int, ...] (#917).
tuple[float|int, ...] (#917). (57d51bb)pdfplumber.repair(...) and .open(repair=True) (#824). (db6ae97)quantize=True, colors=256, bits=8 arguments/defaults to PageImage.save(...). (b049373)Make word segmentation (via WordExtractor.char_begins_new_word(...)) more explict and rigorous; should help in catching edge-cases in the future. (6ac
WordExtractor.char_begins_new_word(...)) more explict and rigorous; should help in catching edge-cases in the future. (6acd580 + ebb93ea + #840)curve_edge objects (instead of just line and rect_edge objects) in default table-detection strategy. (6f6b465 + #858)ffi to ffi), and add the expand_ligatures boolean parameter to text-extraction methods. (86e935d + #598)Page.extract_text_lines(...) method. (4b37397 + #852)main_group, return_groups, return_chars parameters to Page.search(...). (4b37397).curve_edges property to PDF and Page. (6f6b465)Fix x0>x1/etc. error for when drawing rect fills, per new Pillow version
x0>x1/etc. error for when drawing rect fills, per new Pillow version (db136b7)Breaking change: In Page.extract_table[s](...), keep_blank_chars must now be passed as text_keep_blank_chars, for consistency's sake.
Page/utils.extract_text(layout=True) approach so that it pads, to the degree necessary, the ends of lines with spaces and the end of the text with blank lines to acheive better mimicry of page layout. (d3662de)pts attribute and, in doing so, deprecate the curve_obj["points"] attribute, and fix PageImage.draw_line(...)'s handling of diagonal lines. (216bedd)Page.extract_table[s](...), keep_blank_chars must now be passed as text_keep_blank_chars, for consistency's sake. (c4e1b29)Page.extract_table[s](...) support for all Page.extract_text(...) keyword arguments. (c4e1b29)height and width keyword arguemnts to Page.to_image(...). (#798 + 93f7dbd)layout_width, layout_width_chars, layout_height, and layout_width_chars parameters to Page/utils.extract_text(layout=True). (d3662de)None. (#811) [h/t @toshi1127]utils.py into utils/ submodules. Retains same interface, just an improvement in organization. (6351d97)utils.extract_text(...), via Page.extract_text(...), via Page.extract_table(...)). (3424b57)Bump pinned pdfminer.six version to 20221105.
pdfminer.six version to 20221105. (e63a038)text attribute to .textboxhorizontal/etc., regression introduced in 9587cc7 / v0.6.2. (8a0c126)lru_cache usage, which are discouraged for class methods due to garbage-collection issues. (e3142a0)nbexec development requirement from 0.1.0 to 0.2.0. (30dac25)…tuple, etc.) as the key_fn parameter, reverting breaking change in 58b1ab1. (#691 + 1e97656) [h/t @jfuruness]
py.typed file was not included in PyPi distribution. (#698 + #703 + 6908487) [h/t @jhonatan-lopes]utils.cluster_objects(...) with any hashable value (str, int, tuple, etc.) as the key_fn parameter, reverting breaking change in 58b1ab1. (#691 + 1e97656) [h/t @jfuruness]Add utils.outside_bbox(...) and Page.outside_bbox(...) method, which are the inverse of utils.within_bbox(...) and Page.within_bbox(...). (#369 + 3ab1
utils.outside_bbox(...) and Page.outside_bbox(...) method, which are the inverse of utils.within_bbox(...) and Page.within_bbox(...). (#369 + 3ab1cc4)strict=True/False parameter to Page.crop(...), Page.within_bbox(...), and Page.outside_bbox(...); default is True, while False bypasses the test_proposed_bbox(...) check. (#421 + 71ad60f).to_image(...) raises PIL.Image.DecompressionBombError. (#413 + b6ff9e8)PageImage conversions for PDFs with cmyk colorspaces; convert them to rgb earlier in the process. (28330da)Quick fix for transparency issue in visual debugging mode. b98dd7c
Add split_at_punctuation parameter to .extract_words(...) and .extract_text(...). (#682) [h/t @lolipopshock]
split_at_punctuation parameter to .extract_words(...) and .extract_text(...). (#682) [h/t @lolipopshock].to_image(...)'s approach, preferring to composite with a white background instead of removing the alpha channel. (1cd1f9a)LayoutEngine.calculate(...) when processing char objects with len>1 representations, such as ligatures. (#683)Fix bug when calling PageImage.debug_tablefinder() (i.e., with no parameters). (#659 + 063e2ed) [h/t @rneumann7]
Add "matrix" property to char objects, representing the current transformation matrix.
"matrix" property to char objects, representing the current transformation matrix. (ae6f99e)pdfplumber.ctm submodule with class CTM, to calculate scale, skew, and translation of a current transformation matrix obtained from a char's "matrix" property. (ae6f99e)page.search(...), an experimental feature that allows you to search a page's text via regular expressions and non-regex strings, returning the text, any regex matches, the bounding box coordinates, and the char objects themselves. (#201 + 58b1ab1)--include-attrs/--exclude-attrs to CLI (and corresponding params to .to_json(...), .to_csv(...), and Serializer. (4deac25)py.typed for PEP561 compatibility and detection of typing hints by mypy. (ca795d1) [h/t @jhonatan-lopes]pdfminer.six version to 20220524. (486cea8)utils.collate_chars(...), the old name (and then alias) for utils.extract_text(...). (24f3532)The main news about this version is that it introduces __type annotations__, and enforces them via mypy --strict. It also fills in the few remaining g
The main news about this version is that it introduces type annotations, and enforces them via mypy --strict. It also fills in the few remaining gaps in the library's test coverage (although all parts of the library could still use stronger tests). See CHANGELOG.md for details.
mypy --strict. (cdfdb87)TableSettings class, a behind-the-scenes handler for managing and validating table-extraction settings. (9587cc7).to_csv(...) and .to_json(...) from types to object_types. (9587cc7).to_json(...) so that, if an object type is not present for a given page, it has no key in the page's object representation. (9587cc7)utils.filter_objects(...) and move the functionality to within the FilteredPage.objects property calculation, the only part of the library that used it. (9587cc7)pdfminer.pdftypes.STRICT = True and pdfminer.pdfinterp.STRICT = True, since that has now been the default for a while. (9587cc7)See CHANGELOG.md for details. Summary:
See CHANGELOG.md for details. Summary:
pdfminer.six version to 20220319pdfminer.six version to 20220319. (e434ed0)Pillow version to >=9.1. (d88eff1)See CHANGELOG.md for a full list of additions, changes, and fixes. In some (hopefully) rare cases, this version may introduce breaking changes, which…
See CHANGELOG.md for a full list of additions, changes, and fixes. In some (hopefully) rare cases, this version may introduce breaking changes, which is why we're bumping to v0.6.0. Highlights from the changelog include:
pdfminer.six from 20200517 to 20211012; see that library's changelog for details, but a key difference is an improvement in how it assigns line, rect, and curve objects. (Diagonal two-point lines, for instance, are now line objects instead of curve objects.) (#515).extract_text(layout=True), an experimental feature which attempts to mimic the structural layout of the text on the page. (#10)pdfminer.six (#346 + #520).extract_text(...) returns "" instead of None when character list is empty. (#482 + cb9900b) [h/t @tungph]--precision argument to CLI (#520)snap_x_tolerance and snap_y_tolerance to table extraction settings. (#51 + #475) [h/t @dustindall]join_x_tolerance and join_y_tolerance to table extraction settings. (cbb34ce).extract_words(...) now includes doctop among the attributes it returns for each word. (66fef89)And many thanks to @samkit-jain for his feedback and review of contributions to this release. 🎉
.extract_text(layout=True), an experimental feature which attempts to mimic the structural layout of the text on the page. (#10)utils.merge_bboxes(bboxes), which returns the smallest bounding box that contains all bounding boxes in the bboxes argument. (f8d5e70)--precision argument to CLI (#520)snap_x_tolerance and snap_y_tolerance to table extraction settings. (#51 + #475) [h/t @dustindall]join_x_tolerance and join_y_tolerance to table extraction settings. (cbb34ce)pdfminer.six from 20200517 to 20211012; see that library's changelog for details, but a key difference is an improvement in how it assigns line, rect, and curve objects. (Diagonal two-point lines, for instance, are now line objects instead of curve objects.) (#515)pdfminer.six (#346 + #520).extract_text(...) returns "" instead of None when character list is empty. (#482 + cb9900b) [h/t @tungph].extract_words(...) now includes doctop among the attributes it returns for each word. (66fef89)text_strategy, so that it uses the top and bottom of every word, not just the top of every word and the bottom of the last. (#467 + #466 + #265) [h/t @bobluda + @samkit-jain]table.merge_edges(...) behavior when join_tolerance (and x/y variants) <= 0, so that joining is attempted regardless, to handle cases of overlapping lines. (cbb34ce).extract_words(...)/WordExtractor.iter_chars_to_words(...) on very long words, caused by repeatedly re-calculating bounding box. (#483)UnicodeDecodeError when trying to decode utf-16-encoded annotations (#463) [h/t @tungph](text|intersection)_(x|y)_tolerance settings. (#539) [h/t @yoavxyoav]pdfplumber.load(...) method, which has been deprecated since 0.5.23 (54cbbc5)Change .convert_csv(...) to order objects first by page number, rather than object type.
From CHANGELOG.md:
--laparams flag to CLI. (#407).convert_csv(...) to order objects first by page number, rather than object type. (#407).convert_csv(...), .convert_json(...), and CLI so that, by default, they returning all available object types, rather than those in a predefined default list. (#407).extract_text(...) so that it can accept generator objects as its main parameter. (#385) [h/t @alexreg]LTAnno objects (which have no bounding-box coordinates) are not extracted. (Was only an issue when setting laparams.) (#388)Page.extract_table(...) so that it honors text tolerance settings (#415) [h/t @trifling]Fix regression (introduced in 0.5.26/b1849f4) in closing files opened by PDF.open
From CHANGELOG.md:
0.5.26/b1849f4) in closing files opened by PDF.opentextboxhorizontal) when laparams is passed to pdfplumber.open(...). Had been removed in 0.5.24 via 1f87898. (#359 + #364)python setup.py build sdist test to main GitHub action. (#365)Add Page.close/__enter__/__exit__ methods, by generalizing that behavior through the Container class
Page.close/__enter__/__exit__ methods, by generalizing that behavior through the Container class (b1849f4)Decimal objects and do not round themTableFinder to return tables in order of topmost-and-then-leftmost, rather than leftmost-and-then-topmost (#336)Page.to_image()'s handling of alpha layer, to remove aliasing artifacts (#340) [h/t @arlyon]psf/black and flake8 on tests/ (#327Add new boolean argument strict_metadata (default False) to pdfplumber.open(...) method for handling metadata resolution failures
strict_metadata (default False) to pdfplumber.open(...) method for handling metadata resolution failures (f2c510d)setup.py (7854328) (#304)pdfplumber.open(...) so that it does not close file objects passed to it (408605f) (#312)Added extra_attrs=[...] parameter to .extract_text(...)
extra_attrs=[...] parameter to .extract_text(...) (c8b200e) (#28)utils/page.dedupe_chars(...) (04fd56a + b132d45) (#71)upright from int to bool (per original pdfminer.six representation) (1f87898)Container.figures, given that they are not fundamental objects (8e74cb9)explicit_horizontal_lines/explicit_vertical_lines descs passed to TableFinder methods (bc40779) (#290)See changelog for details.
See changelog for details.
utils.resolve (non-recursive .resolve_all) (7a90630)page.annots and page.hyperlinks, replacing non-functional page.annos, and mirroring pdfminer's language ("annot" vs. "anno"). (aa03961)page/pdf.to_json and page/pdf.to_csv (cbc91c6)relative=True/False parameter to .crop and .within_bbox; those methods also now raise exceptions for invalid and out-of-page bounding boxes. (047ad34) [h/t @samkit-jain]pdfminer.from_path and pdfminer.load as deprecated; now pdfminer.open is the canonical way to load a PDF. (00e789b).extract_words, which had been returning incorrect results when horizontal_ltr = False (d16aa13)utils.resize_object, which had been failing in various permutations (d16aa13)lines_strict table-finding strategy, which a typo had prevented from being usable (f0c9b85)utils.resolve_all to guard against two known sources of infinite recursion (cbc91c6)pandas from dev requirements and tests (a5e7d7f)Upgraded pdfminer.six requirement to ==20200517 (cddbff7) [h/t @youngquan]
pdfminer.six requirement to ==20200517 (cddbff7) [h/t @youngquan]non_stroking_color attribute on char objects (0254da3) [h/t @idan-david]Fix Page.extract_table(...) to return None instead of crashing when no table is found (d64afa8) [h/t @stucka]
Page.extract_table(...) to return None instead of crashing when no table is found (d64afa8) [h/t @stucka]Fix .get_page_image to prefer paths over streams, when possible (ab957de) [h/t @ubmarco]
Add utils.decimalize performance improvement (830d117) [h/t @ubmarco]
Allow rect and curve objects also to be passed to "explicit_..._lines" setting when table-finding. (And disallow other types of dicts to be passed.)
rect and curve objects also to be passed to "explicit_..._lines" setting when table-finding. (And disallow other types of dicts to be passed.)utils.extract_text bug introduced in prior versionFix and simplify obj-in-bbox logic (see commit 25672961)
utils.extract_text handles vertical text (see commit 8a5d858b) [h/t @dwalton76]Page.to_image use bytes stream instead of file path (Issue #124 / PR #179) [h/t @cheungpat]Page.extract_tables did not pass kwargs to Table.extract [h/t @jsfenfen]Prevent custom LAParams from raising exception (Issue #168 / PR #169) [h/t @frascuchon]
Primarily: Upgrades pinned requirements for pdfminer.six and pillow.
Primarily: Upgrades pinned requirements for pdfminer.six and pillow.
pdfminer.six requirement to ==20200104pillow requirement >=7.0.0tox testsFix sorting bug in page.extract_table()
page.extract_table()Fixed PDF object resolution for rotation (PR #136)
Support for password-protected PDFs
cdecimal support for Python 2Caching for .decimalize() method
.decimalize() methodpdfminer.six==20181108Fix bug in which, when calling get_page_image(...), the alpha channel could make the whole page black out.
Fix issue #67, in which bool-type metadata were handled incorrectly
Fix issue #53, in which non-decimalize-able (non_)stroking_color properties were raising errors.
.travis.yml, but failing on .to_image()
.travis.yml, but failing on .to_image()pycrypto to pycryptodomepdfminer.six to 20170720Fix issue #41, in which PDF-object-referenced cropboxes/mediaboxes weren't being fully resolved.
Access to __version__ from main namespace
__version__ from main namespacedecode_text's argument typePin pdfminer.six to version 20151013 (for now), fixing incompatibility
pdfminer.six to version 20151013 (for now), fixing incompatibilityAllow import pdfplumber even if ImageMagick not installed.
import pdfplumber even if ImageMagick not installed.Access to curve points. (E.g., page.curves[0]["points"].)
curve points. (E.g., page.curves[0]["points"].).draw_line to draw curve points.utils.decimalize a bit more robust; now throws errors on non-decimalizable items.pdfminer object attributes..draw_line from a bounding box to ((x, y), (x, y)), for consistency with curve["points"] and with Pillow's underlying method..rect_edges is called before .edgesQuick-draw PageImage methods: .draw_vline, .draw_vlines, .draw_hline, and .draw_hlines.
PageImage methods: .draw_vline, .draw_vlines, .draw_hline, and .draw_hlines.keep_blank_chars for .extract_words(...) and TableFinder settings.text_tolerance and intersection_tolerance TableFinder values from 1 to 3.pillow images.pandas DataFrames as inputs to multi-draw commands (e.g., PageImage.draw_rects(...)).Completely overhauls the approach to table extraction.
Page.to_image(...) and PageImage. (Introduces wand and pillow as package requirements.).crop from .intersects_bbox and .within_bbox.x_tolerance and y_tolerance for word extraction from 5 to 3Provide access to Page.page_number
Page.page_number.page_number instead of .page_id as primary identifier. [h/t @jsfenfen]x_tolerance and y_tolerance for word extraction from 0 to 5Fix bug stemming from when metadata includes a PostScript literal. [h/t @boblannon]
Your coding agent can read these notes before it upgrades. Set up the MCP server →