NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1187 most downloaded on PyPI
Read, write, repair, and transform PDFs in Python, powered by qpdf
Last release 5 days ago
29 Sep 2026
Ships fairly regularly
a new release about every 3 weeks
Some releases are documented
notes for 17 of the last 60 stable releases
4 versions withdrawn
withdrawn after publishing
9 years old
275 releases · first in 2017
Added binary wheels for Windows on ARM64. Thanks to @ndabas . {issue} 744
744pikepdf.sanitize.remove_javascript,pikepdf.sanitize.remove_external_access,pikepdf.sanitize.remove_multimedia and the matchingpikepdf.sanitize.Sanitizer steps missed actions attached to form/Kids, and on widgets that appear on no/AcroForm field tree. Thanks to/Next chain untouched past its/JavaScript action)746/Next chain that is not anAttributeError or ValueError on some such values/Next [42]), and now runs in explicit conversion modeOne column per quarter.
Attaching a file through Pdf.attachments now lists its file specification in the document catalog's /AF (associated files) array, and deleting or repl
Pdf.attachments now lists its file specification/AF (associated files) array, and deleting or/AFRelationship pikepdf already writes, a file attached to a PDF/A-3/EmbeddedFiles463/EmbeddedFiles name tree and the catalog's /AF array are checked against/F and /UF, /AFRelationship,pikepdf.pdfa.prepare gives PDF/A-3 embedded files a missing MIME typeapplication/octet-stream) and /AFRelationship (/Unspecified), and/AF array; newpikepdf.pdfa.PrepareResult fields count these repairs. Seepdfa-attachments./Dests name tree/Dests dictionary, instead of reporting them as/Names /AP) as Form XObjects./AlternatePresentations is now reported as a violation of PDF/A-2 and/Names /JavaScript)pikepdf.pdfa.prepare removes them.pikepdf.pdfa.prepare gives the encoding of a non-symbolic TrueType/Differences but no /BaseEncoding the base encoding/WinAnsiEncoding, as PDF/A-2 and PDF/A-3 require, when no glyph thepikepdf.pdfa.PrepareResult fields report these/Extensions << /ADBE ... >>) that Acrobat writes in the document catalog,pikepdf.pdfa.resolve_save_kwargs accepts an extension level inmin_version or force_version for PDF/A-2 and PDF/A-3, such as('1.7', 8).pikepdf.models.PdfMetadata.copy_properties copies named XMPrdf:Description whatever its rdf:about, and the structure of knownpikepdf.pdfa.repair_annotation_flags, the step ofpikepdf.pdfa.prepare that removes hidden annotations and sets thepikepdf.pdfa.AnnotationRepairResult. Running it before handing a file/DL (decoded length)/Differences only from the Adobe Glyph List for New Fonts, so it rejectedafii10017 for Cyrillic text. It now accepts every nameuni0410 or A.sc, are still rejected.Pdf opened or created with conversion_mode='explicit' returned plainint for some numbers that qpdf writes itself, such as /Pages /Count/Rotate after {meth}pikepdf.Page.rotate,/BBox of {meth}pikepdf.Page.as_form_xobject. These values nowInteger and Real now support the rest ofint or Decimalround(), math.trunc(), math.floor(),math.ceil(), divmod(), and format specs such as f'{x:.2f}'.int() of a Real now truncates, as int() of a Decimal does, insteadTypeError; for example int(page.Rotate) now works on a/Rotate 90.0.Matrix now accept a pikepdf.Object, such as a/Matrix array, and a list of objects, as the runtime always has.macOS wheels are now built on GitHub's macos-15 runner, since GitHub is deprecating the macos-14 runner. Binary wheels now require macOS 15 or newer.
Several improvements to explicit conversion mode and NamePath, prompted by
the OCRmyPDF project's migration to these APIs
(pikepdf.explicit_conversion(), the as_* safe accessors, and NamePath)
in a production codebase that reads untrusted, often malformed, PDFs.
tomllib, enum.StrEnum,typing.Self, datetime.UTC and the broader ISO 8601 support indatetime.fromisoformat. The compatibility shims for Python 3.10 are gone.cp314-abi3 wheel; free-threaded CPython 3.15 gets its owncp315t wheel, since the stable ABI does not cover free-threaded builds.pikepdf.Pdf.open and {meth}pikepdf.Pdf.new now accept aconversion_mode argument ('implicit' or 'explicit'), andpikepdf.Pdf.conversion_mode property. A Pdf's mode travels withNone reverts to inheriting the context-manager/globalpikepdf.implicit_conversion, a thread-local context managerpikepdf.explicit_conversion, so implicit mode can bePdfpikepdf.set_object_conversion_mode. An object with no owning Pdfpikepdf.Dictionary(...) you have not yet attached to aPdf it follows that document's mode.pikepdf.Object.get_raw, which behaves like~pikepdf.Object.get (key, Name, or NamePath, plus a default)pikepdf.ObjectNull-typed object rather than None; in a dictionary,get_raw returns theget does.~pikepdf.Object.get_int,~pikepdf.Object.get_bool, {meth}~pikepdf.Object.get_float,~pikepdf.Object.get_decimal, {meth}~pikepdf.Object.get_dict, and~pikepdf.Object.get_list, each (key_or_path, default=None). Theseget_raw with the matching as_* accessor, so reading an optionalwith pikepdf.explicit_conversion(): ... wrapper.coerce=True to {meth}~pikepdf.Object.as_int,~pikepdf.Object.as_bool, {meth}~pikepdf.Object.as_float, and~pikepdf.Object.as_decimal, and to the corresponding get_*as_int(coerce=True) accepts a Real (truncated toward zero) and aString; as_bool(coerce=True) accepts a nonzero Integer orReal (fixing the common /Marked 1 case that previously required aas_int() != 0); as_float/as_decimal with coerce=TrueInteger and a numeric String, including exponent notation"1e-5".pikepdf.Integer and pikepdf.Real now support the ordering comparisons<, <=, > and >= against Python int, float, bool, Decimal,box[0] < box[2], sorted(), min() and max() workReal compares by its exact decimal value.pikepdf.Integer and pikepdf.Real now works between twoReal('2.5') + Integer(3)) and with Decimal and bool** is supported. Results are the native Python type thatint for Integer with int/Integer,Decimal whenever a Real or Decimal is involved, and float with afloat operand. In explicit mode, box[2] - box[0] > 100 therefore worksTypeError if the PDF stored a non-number.pikepdf.unbox, which returns the native Python value of anInteger, Boolean or Real and passes any other value throughas_* accessors it does not require knowing whichisinstance, is True,Decimal(), json.dumps or a function documented to return a native type.pikepdf.as_int,~pikepdf.as_bool, {func}~pikepdf.as_float,~pikepdf.as_decimal, {func}~pikepdf.as_dict,~pikepdf.as_list, {func}~pikepdf.as_str and~pikepdf.as_bytes, each (value, default=None), with keyword-onlycoerce for the numeric and boolean ones. They are the value-side twins ofget_* typed getters, for an array element, content stream operand oras_* method does, and a native value is accepted only if it is the typepikepdf.as_int(True) gives the default,~pikepdf.Object.as_str and {meth}~pikepdf.Object.as_bytes,String's decoded text or raw bytes, and raise TypeErrorstr() andbytes(), which accept almost anything. Added the matching typed getters~pikepdf.Object.get_str and {meth}~pikepdf.Object.get_bytes./topics/objects for the full description of scopes and/topics/type_safety, explaining whypath in obj now tests whether the path can be traversed to a value, anddel obj[path] deletes the final component after traversing to itsNamePath is now iterable, yielding its individual components (str forint for indices) in order, e.g. list(NamePath.A.B[0]) gives['/A', '/B', 0].NamePath instances now compare and hash by their component sequence, soNamePath.A.B == NamePath.A.B is True and a NamePath can be used as aset.isinstance(path, pikepdf.NamePath) now works, and pikepdf.NamePath candef f(key: pikepdf.Name | pikepdf.NamePath)) without importing a privateOCRmyPDF tested its PDF/A output against veraPDF, and found files that
pikepdf's metadata handling made invalid.
dc:title and dc:description are nowx-default item, as the XMP specification defines, ratherx-default last when a document/Title and /Subject were synced from the wrongx-default item and keeps the other languages, instead of replacing all ofx-default item, instead of writing several x-default items.rdf:Description, not onlyrdf:about="". Properties in Descriptions withrdf:about="uuid:..." (written by Distiller and older Acrobat) or with nordf:about could not be read or deleted, and setting one added a duplicate,rdf:Descriptionrdf:about value, as the XMP specification requires, andpikepdf.models.PdfMetadata.recovered andXmpDocument.recovered, which are True if the XMP was not well-formed andpikepdf.PdfInlineImage.read_raw_bytes, which returns thepikepdf.Pdf.save now corrects /Count in each node of the page/Count as it was, so a wrong /Count in the input was writtenpikepdf.pdfa module, which prepares, checks and savespdfa.
pikepdf.pdfa.prepare repairs a document in memory: it installs a/CIDSet for PDF/A-1,~pikepdf.pdfa.PrepareResult describing what changed:~pikepdf.pdfa.PrepareResult.describe gives a sentence per change,~pikepdf.pdfa.PrepareResult.messages the same sentences with aPrepareResult.xmp_problem why an XMP packetrdf:Descriptionrdf:about, including the non-empty uuid:... valuepikepdf.pdfa.check predicts the verdict for the file that would bepikepdf.pdfa.save prepares the document, writes it to a temporary~pikepdf.pdfa.PdfaError with the report (whose prepared recordspikepdf.pdfa.resolve_save_kwargs returns the completepikepdf.Pdf.save settings for a flavour. Settings that PDF/Aencryption, normalize_content and fix_metadata_version) are pinned,ValueError.~pikepdf.pdfa.Report has a verdict of 'pass','fail' (at least one violation) or 'not_checked' (only constructs thesave accepts only'pass'. Each {class}~pikepdf.pdfa.Finding names a rule: a veraPDF rule idISO_19005_2:6.2.8-3, or a pikepdf: id for local policies andprepare andcheckpikepdf.pdfa needs jsonschema, referencing and fontTools, available aspip install 'pikepdf[pdfa]'. import pikepdf does notimport pikepdf.pdfa raises ImportError naming the extra ifthird-party-licenses/README.md.pikepdf._io.atomic_write_verified, which writes apikepdf.pdfa.save uses it.pikepdf.settings.get_qpdf_limits andpikepdf.settings.set_qpdf_limits, which read and change qpdf'spikepdf.JobBuilder.limits. set_qpdf_limitspikepdf.settings.disable_qpdf_default_limits, which liftspikepdf.settings.qpdf_limit_errors, which counts how many times anypikepdf.parse_content_stream,pikepdf.Page.parse_contents) now stops after 15 syntax errors,pikepdf.settings.set_qpdf_limits(parser_max_errors=0) to parse as muchpikepdf.PdfError when the pages are accessed./Rotate values outside [0, 360), such as -90, are now normalizedprogname keyword argument of {class}pikepdf.Job is deprecated andDeprecationWarning. It was passed toQPDFJob::initializeFromArgv as the name of an environment variable,args, as it always was.PdfImage metadata such as .width,.height or .colorspace whose value has the wrong PDF type (for example/Width written as a string or name) now raises TypeError, with aNotImplementedError: Metadata access for Width. The exception is also aNotImplementedError, so existing handlers keep working.pikepdf.String no longer compares equal tobytes: pikepdf.String('abc') == b'abc' is now False. It stillstr it decodes to. A String compared equal tostr and a bytes that are unequal to each other, and so could not{pikepdf.String('héllo'): 1}['héllo'] raised KeyError. A String nowstr. To compare the raw data, use bytes(s) == b'...' or~pikepdf.Object.as_bytes.pikepdf.Name no longer compares equal to bytes:pikepdf.Name('/Foo') == b'/Foo' is now False. It still compares equal tostr of its UTF-8 bytes, so obj.Type == '/Page' works as before, andstr, so {pikepdf.Name('/héllo'): 1}['/héllo']KeyError. As PDF 2.0 (ISO 32000-2, 7.3.5) requires,bytes(name) == b'...' to compare the raw bytes.pikepdf.Real combined with an int, Integer, orDecimal, or negated with unary -/+/abs(), now yields a Decimalfloat, or TypeError for int operands other than /).Real with a float operand still yields a float. Division of anInteger or Real by zero now raises ZeroDivisionError, as for PythonValueError.pikepdf.Matrix now raises TypeError, notValueError, for a pikepdf.Object that is not a matrix or anObjectList with a non-numeric element, matching what a native argument ofexcept TypeError around Matrix(*operands)must have 6 elements) remain ValueError.bool() on a pikepdf.Integer or pikepdf.Real isbool(pikepdf.Integer(0)) is False), instead of raisingNotImplementedError: code is unreachable.~pikepdf.Object.as_dict and~pikepdf.Object.as_list now accept a default argument and raiseTypeError on a type mismatch, instead of raising pikepdf.PdfError withPdfpdf.Root.X = 42 followed bypdf.Root.get_raw('/X').is_owned_by(pdf) is True, matching objects parsedName, String, Integer, ...) are adopted as a copy,Dictionary or Array is adopted in place and keeps its aliasing with thePdf now raisesForeignObjectError, as it already did for containers parsed from a file.pikepdf.Pdf.copy_foreign. Adoption isPdf. If the same object was reachable under two keys, removingpikepdf.Pdf.copy_foreign, {meth}pikepdf.Object.with_same_owner_as,pikepdf.Pdf.make_indirect, or by appending/inserting a page intoPdf.pages, and values inserted through NameTree/NumberTree__setitem__, now report the destination document as their owner and followconversion_mode.pikepdf.Object.with_same_owner_as andpikepdf.Pdf.make_indirect now raise ForeignObjectError for aPdf, instead of silentlypikepdf.Pdf.copy_foreign.~pikepdf.Object.as_int with coerce=True nowOverflowError. Called without a default itOverflowError.repr() of a pikepdf.Object now honors thePdf, then the global setting), rather than only thepikepdf.set_object_conversion_mode now raisesValueError for a value other than 'implicit' or 'explicit', insteadpikepdf.Pdf.save writes directly to the file descriptor when savingopen, instead ofwrite() method for every chunk of output. Saving a large document is about twice as fast. Other streams, suchBytesIO, pipes and subclasses of the io classes, are written throughwrite() as before.isinstance() checks against {class}pikepdf.Dictionary,pikepdf.Stream, {class}pikepdf.Integer and the other objectpikepdf.unbox is now implemented in C++ and is about 40 timespikepdf.as_int, {func}pikepdf.as_float andpikepdf.as_decimal answer for a native int or Decimal withoutDecimal no longerdecimal module each time.pikepdf.models.metadata.decode_pdf_date now accepts every PDF dateD:202001011230 (no seconds) was misread as 12:03, andD:2020010112 (no minutes) and offsets in hours only, such as -08',ValueError, so {attr}pikepdf.Pdf.docinfo dates in these formspikepdf.Pdf.save now raises ValueError when the destination isPdf was opened from, or another stream on the same file,Pdf, since qpdf reads its inputallow_overwriting_input=True to save over the input.<dc:title>Old</dc:title>, now replaces the text. Previously the new valuepikepdf.PdfImage.filter_decodeparms andpikepdf.PdfImage.decode_parms no longer raise AttributeError when/DecodeParms is a bare number or boolean, in either conversion/DecodeParms that is neither a dictionary nor an array is now/Decode that is not an array (a bare number, boolean, name orAttributeError when the image is read or converted.pikepdf.Stream.write now raises TypeError rather thanAttributeError when filter or decode_parms is a native Python valueint, str or dict.pikepdf.Pdf.save to an existing file that is not a regular file,/dev/null, a FIFO or a character device, now writes into it/dev/null replaced the device with a regular file.pikepdf.Matrix.inverse now returns the correct translation fore and f terms wereint outside the signed 64-bit range of a PDF integer, byArray(...), Dictionary(...) or Integer(...), nowOverflowError instead of RuntimeError: std::bad_cast (or aTypeError).PdfImage, PdfImageBase, PdfJpxImage, PdfInlineImage andPaletteData once again report their module as pikepdf.models.image, thepikepdf.models.image._classes introduced when that module became a packagepikepdf.set_object_conversion_mode andpikepdf.explicit_conversion no longer open a nonexistent test.pdf,str() of a pikepdf.Integer, Boolean or Real now gives the value --'42', 'True', '1.50' -- rather than the object's repr. In implicitint/bool/Decimal and str() never reachedhash() now works on pikepdf.Integer, Boolean and Real (and on aRuntimeError: don't know how to hash this. A scalar hashes like the Python value it compares equal to, soReal('1.0'), Decimal('1.0') and 1 agree, and a scalar can be a dictpikepdf.Pdf.make_indirect and {meth}pikepdf.Object.with_same_owner_asName, String, Integer, Real,Boolean, Operator) that you pass them into an indirect object in place.Pdf, after which hash() raised and any dict or set alreadyFOO = Name.Foo becamerepr() of an Array, Dictionary or Stream nowpikepdf.Real('42.42') rather than'42.42', which was indistinguishable from a PDF string and did noteval(repr(obj)).PdfImage raised NotImplementedError for any indexed or /DeviceNpage.label restarted numbering at 1 and warned/St), outlines (a closed item read as open, flags came/Count), form appearanceTypeError laying out multiline and combed text fields),get_objects_with_ctm (a malformed cm operator raised instead of beingSimpleFont metrics, and Action.new_window.PdfImage metadata now reads the same in either conversion mode. A real/Width or /Height or a boolean /BitsPerComponent raisedNotImplementedError in explicit mode, and a real inside a /ColorSpaceSimpleFont.ascent, descent and unscaled_char_width() now return aDecimal as documented, in either conversion mode. Font metrics stored asint.NamePath (e.g. list(path)) no longer falls back to the__getitem__(0), (1), ... protocol, which never raisedIndexError and so iterated forever, exhausting memory.NamePath in obj no longer silently answers False for every path; seeNamePath['/A']['/B'].C[0] example in the type stub, which__getitem__ only accepts an int). The correct form, matching the/topics/namepath documentation, isNamePath['/A']('/B').C[0].box = page.obj.get_raw('/MediaBox') followed bydel page.obj['/MediaBox']) and used after the document was closed, kept arepr() or when inserted elsewhere. Removing an object from a document nowPdf disconnects every object itactions/cache, so jobs that share an image no longer each download andsrc/pikepdf/_core.pyi, are a stub-only package src/pikepdf/_core/ split_core/_matrix.pyi coverssrc/core/matrix.cpp, _core/_page.pyi covers src/core/page.cpp, and so_core/__init__.pyi re-exporting the lot and documenting thepikepdf._coreimport pikepdf._core still resolves topikepdf._core._object.Object rather than pikepdf._core.Object);pikepdf.Object remains the name to write in annotations.src/pikepdf/_core/ stubs, which are what Sphinx, type checkers and IDEshelp() on a C++ method now shows only itsRestored the full {class}pikepdf.Stream and {class}pikepdf.Dictionary
documentation, including constructor arguments and examples, which had been
reduced to a single line when those classes moved to C++.
Clarified guidance on Pdf.save(..., deterministic_id=) and static_id=, and
the matching JobBuilder.deterministic_id() and JobBuilder.static_id().
deterministic_id gives reproducible, production-safe /ID values;
static_id sets the same dummy /ID in every file and is for testing only.
static_id now appears last in the Pdf.save() signature. All save()
options are keyword-only, so existing code is unaffected.
A warning raised from pikepdf's C++ layer -- PageCopyWarning , the Page.rotate() deprecation warning, and the several warnings issued while opening an…
pikepdf's exceptions now form a documented hierarchy rooted at the new
pikepdf.PikepdfError, with pikepdf.PikepdfWarning playing the same role for
warnings. See {doc}/api/exceptions for the full tree. {issue}739
pikepdf.DataDecodingError now derives frompikepdf.PdfError. A stream that will not decode is a defect in the document,Object.read_bytes() -- already raisedPdfError for other kinds of damage, so except PdfError was a handler thatDataDecodingError by name is unaffected. Code thatexcept PdfError before except DataDecodingError will now take thepikepdf.PdfParsingError now derives frompikepdf.PdfError, for the same reason.pikepdf.PasswordError remains a sibling of PdfError, not a subclass. Aexcept PdfError not catching it.pikepdf.NotExtractableError is now exported. It was already the base classHifiPrintImageNotTranscodableError but could not be caughtFormCopyWarning (the class is PageCopyWarning) and omittedReferenceCycleError, PageCopyWarning and NotExtractableError.XMP assigns a type to every standard property, and software that reads XMP
discards a property whose type is wrong -- Ghostscript strips these silently,
and PDF/A validators reject them. pikepdf used to choose the RDF container from
the Python type of the value it was handed, so meta['dc:subject'] = ['a', 'b']
wrote an rdf:Seq where the specification requires an rdf:Bag, and
meta['dc:creator'] = 'Author' wrote a bare string where an rdf:Seq belongs.
pikepdf now knows the type of the properties in the standard schemas and
converts values to it. See {ref}metadatatypes. {issue}555
list or aset was assigned. Reading such a property back returns a list for anset for an unordered one, as before.pikepdf.XmpTypeWarning and stores a single elementlog.error message pikepdf previously produced for dc:creator alone islist or set to a property that holds one; (warning that it did so) instead ofdatetime.datetime and datetime.date may now be assigned to date valuedxmp:CreateDate, and are encoded as ISO 8601. A time zone+hh:mm and -hh:mm. int may be assigned to pdfaid:partbool to xmpRights:Marked.set to an ordered property such as dc:creator now sorts thePdf.open_metadata(strict=True) now also raises TypeError for a value thatKeyError namingpikepdf.models.metadata.XmpDocument.register_xml_namespace(). Previously thepikepdf.models.metadata.XMP_SCHEMA maps a property's qualified nameXmpProperty type, for callers that want to check types themselves.xmp:MetadataDate rather{http://ns.adobe.com/xap/1.0/}MetadataDate -- so iteration yielded a__getitem__ resolved to a different one, and any code that walkeddict(meta.items())) raised KeyError or, where theValueError: Invalid tag name. pikepdf now rebinds such634key in metadata now reports whether the key is present rather than whetherin while metadata[key] returns its value. Iterating therdf:Description.datetime.fromisoformat. On Python 3.10 the2024-06-01T12:00:00.5Z in XMP caused /ModDate to be dropped from2024-06-01 nowD:20240601 in DocumentInfo, and back, without a spurious midnight'xmp:' now raises KeyError (andmetadata.get() returns the default) instead of ValueError from lxml.pikepdf._core._ObjectList, the list of operands attached to a content streampikepdf.Object, but the elements of an operand listint/bool/Decimal on the==, !=,in, count(), remove(), append(), insert(), extend() and__setitem__ now encode their argument the same way pikepdf.Array does, soinstruction.operands == [0] is True and 0 in instruction.operands works.742_ObjectList now compare objects by value, as the rest ofoperator==, which reports only==, in,count() and remove() gave wrong answers even for operand lists madepikepdf.Object.pikepdf._core._ObjectMapping, returned by Object.as_dict() andPage.get_images(), had all of the same problems and received the same__setitem__ and update() now encode their value, == and !=dict works instead of printing a nanobind conversion warning._ObjectMapping.__setitem__ and __delitem__ now accept a pikepdf.Name__getitem__ and __contains__ already did._ObjectList or _ObjectMapping to a list or dict no longernanobind: implicit conversion from type 'list' to type 'pikepdf._core._ObjectList' failed! to stderr.PageCopyWarning, thePage.rotate() deprecation warning, and the several warnings issued whilepython -W error orwarnings.simplefilter('error'). Previously the exception was created andrepr() of a deeply nested object and unparse_content_stream() no longerpikepdf.StreamParser is now exported from the top-level package and included__all__. It was always the required argument type of the publicPage.parse_contents(), but previously could only be imported from thepikepdf._core module. {issue}738Page's box properties (mediabox, cropbox, artbox, bleedbox,trimbox), the rotation property and rotate(), and _ObjectMapping'sget, __getitem__, __setitem__, __delitem__,__contains__) from Python augmentations to C++. Each was a thin PythonPage._get_mediabox(), _get_artbox(), _get_bleedbox(), _get_cropbox(),_get_trimbox() and _get_rotation() bindings they delegated to are gone.augment_override_cpp decorator from pikepdf._augments. A_cpp<name> copy the decoratorPage._cpp__repr__ no longer exist.Attachments mapping methods, AttachedFileSpec.relationship and its__repr__, AttachedFile.read_bytes(), Page.form_xobjects, Rectangle's__repr__, __hash__ and to_bbox(), Token.__repr__, and Object'sas_int(), as_bool(), as_float(), as_decimal() and_ipython_key_completions_(). len(pdf.attachments) and iterating it nopikepdf.AttachedFileSpec for every attached file merely toAttachments._get_all_filespecs(), _get_filespec(), _attach_data(),_add_replace_filespec(), _remove_filespec(), Page._form_xobjects andObject._get_real_value() -- are gone.augments may now subclass an abstract base classpikepdf.Attachments uses this to get the MutableMapping mixins.Nothing published for this version
Native crypto implements MD5, RC4, SHA2 and AES itself, so no upstream deprecation policy can withdraw the weak algorithms older PDFs require. There i…
Binary wheels now redistribute the licenses of the compiled third-party
libraries they bundle, along with an attribution manifest mapping each
component to its license. {issue}736
third-party-licenses/ directory documents every vendored binary:qpdf30.dll) plus the Microsoft Visual C++ runtime on Windows;project.license-files, so they shippikepdf-<version>.dist-info/licenses/ and are enumerated in the wheel'sLicense-File metadata. License-Expression remains MPL-2.0: pikepdf'slicenses-for-wheels.txt, which has been removed. It hadSource distributions no longer contain qpdf's source tree. CI unpacks qpdf
into ./qpdf to build libqpdf, and that build prerequisite was being swept
into the sdist, adding roughly 2900 Apache-2.0 files to a distribution that
declares itself MPL-2.0. sdists now contain only pikepdf's own code.
macOS wheels now use qpdf's native crypto provider instead of GnuTLS, and all
POSIX builds select it explicitly rather than relying on qpdf's implicit
fallback. macOS wheels lose nine bundled libraries as a result -- GnuTLS,
Nettle, Hogweed, GMP, libidn2, libunistring, Libtasn1, p11-kit and libintl --
along with roughly 10 MB and the LGPL obligations that came with them.
macOS moved to GnuTLS in v8.11.1 to fix legacy encrypted files failing to open
({issue}520), because Homebrew's OpenSSL had retired the legacy provider
that supplies the RC4 and MD5 those files need. The Linux builds were assumed
to be doing the same thing, but were not: their images carry no GnuTLS or
OpenSSL headers, so they had silently been using native crypto all along, and
have been opening legacy encrypted files without trouble ever since. Native
crypto implements MD5, RC4, SHA2 and AES itself, so no upstream deprecation
policy can withdraw the weak algorithms older PDFs require. There is no
change to which files pikepdf can open.
Documentation updates with additional guidance about crypto provider selection
and the realities of PDF encryption security.
pikepdf.PdfInlineImage.icc now returns None instead of raising~pikepdf.exceptions.InvalidPdfImageError. An inline image's colour/ICCBased, so "no ICC profile" is the correct answer ratherpikepdf.PdfImageBase.palette206
/F [/AHx /Fl] reports['/ASCIIHexDecode', '/FlateDecode'] frompikepdf.PdfInlineImage.filters, and an /Indexed colour space/CS [/I /RGB 1 <...>] is now recognized as indexed instead ofNotImplementedError. Correspondingly,pikepdf.PdfInlineImage.unparse now abbreviates names inside arrays,/DecodeParms dictionaries are left/I is /Interpolate/Indexed as a value (ISO 32000-2 Tables 91 and 92)./I was always expanded to /Indexed, so /I true became/Indexed true./Fl (/FlateDecode), /D (/Decode)/L (/Length); none of the three were expanded before.~pikepdf.exceptions.DataDecodingError with the decoder's errorRuntimeError("qpdf will consume this exception")~pikepdf.exceptions.DataDecodingError raised by~pikepdf.exceptions.DependencyError from a customKeyboardInterrupt -- was reported as if the JBIG2 data werepikepdf.jbig2.JBIG2DecoderInterface.decode_jbig2 should raiseDataDecodingError to report undecodable data.pikepdf.Pdf.get_warnings rather than in the exception.Extended PDF outline (bookmark) support to cover more of the spec:
pikepdf.OutlineItem now supports the outline item dictionary's/C (color), /F (italic/bold flags via {class}pikepdf.OutlineItemFlag.italic/.bold properties), and /SE (structure elementpikepdf.Destination, which parses an existing destinationpikepdf.make_page_destination.pikepdf.OutlineItem.resolved_destination resolves any form of an/Dests dictionary) -- to aDestination.pikepdf.Action and subclassesGoToAction, GoToRAction, GoToEAction, GoToDpAction, LaunchAction,URIAction, NamedAction, SetOCGStateAction, JavaScriptAction) for/A action dictionaries.pikepdf.OutlineItem.parsed_action gives a typed view of an item's/Count being written as 0 instead/Indexed images losing their palette when extracted.pikepdf.PdfImage.as_pil_image and {meth}pikepdf.PdfImage.extract_to/DeviceCMYKhival 0). Both now decode correctly, matching the existing behavior for/Indexed image whose base colour space is unsupported (for example/DeviceN or /Separation) now raises NotImplementedError instead of/Width and /Height before allocating, raising~pikepdf.exceptions.ImageDecompressionError when it does not. Thispikepdf.PdfImage.MAX_IMAGE_PIXELS limit added in v10.10.0: thatNoneMAX_IMAGE_PIXELS.~pikepdf.exceptions.ImageDecompressionError rather than returning a/Width or /Height raises~pikepdf.exceptions.InvalidPdfImageError instead of failing further{meth} pikepdf.PdfImage.as_pil_image and {meth} pikepdf.PdfImage.extract_to now apply an image's soft mask ( /SMask ) or explicit/colour-key mask ( /M
pikepdf.PdfImage.as_pil_image and {meth}pikepdf.PdfImage.extract_to/SMask) or explicit/colour-key mask/Mask) by default, returning an image with an alpha channel (LA orRGBA) and writing a transparency-capable format (.png). Previouslyapply_mask=False to recover the old/SMask soft masks,/Mask stencil (explicit) masks, and /Mask colour-key masks are/CalRGB, /CalGray and /CalCMYK images, whichNotImplementedError. The samples are decoded as theirWhitePoint/Gamma/Matrix is attached to the/Lab colour space, extracted as a Pillow LABL*/a*/b* ranges remapped toI;16; 16-bit RGB and CMYK are reduced to 8-bit (withNotImplementedError: ICCBased CMYK images and indexed images now produce/Decode array, and inline images that name their/Resources when obtained viapikepdf.parse_content_stream./DCTDecode) images with a non-default /ColorTransform -- a YCCK/DCTDecode,/CCITTFaxDecode, /JPXDecode, /JBIG2Decode) in any number ofpikepdf.PdfImage.MAX_IMAGE_PIXELS, a settable class-level limitmax(500_000_000, PIL.Image.MAX_IMAGE_PIXELS) -- a floorNone to disable the check.pikepdf.Array now implements the standard Python list interface:del on slices and slice assignment), and theclear(), count(), index(), insert(), pop(), remove(),reverse() methods./SMaskInData (alpha encoded inside a JPEG 2000 stream) is applied onlySMaskInData 2) result is not un-premultiplied. A/Matte entry on a soft mask is not undone (a warning is emitted). When an/SMask and a /Mask, the soft mask takes precedence.L/RGB/CMYK images.[/DCTDecode /CCITTFaxDecode]) cannot be decoded by any reader and now~pikepdf.exceptions.UnsupportedImageTypeError rather thanNotImplementedError./Width and/Height so that {meth}pikepdf.PdfImage.as_pil_image attempted topikepdf.PdfImage.MAX_IMAGE_PIXELSPIL.Image.MAX_IMAGE_PIXELS), across every image-decodepikepdf.DecompressionBombError and borderline images emitpikepdf.DecompressionBombWarning (both subclass Pillow's equivalents).[/FlateDecode /CCITTFaxDecode]) now builds its TIFF header from the/CCITTFaxDecode filter's own /DecodeParms rather than the leading filter's,pikepdf.StreamDecodeLevel: thespecialized and all levels were each described with the other's behavior.Fixed a crash ( SIGABRT via std::terminate ) that could occur when a file-backed {class} pikepdf.Pdf was deallocated while a Python exception was alre
SIGABRT via std::terminate) that could occur when apikepdf.Pdf was deallocated while a Python exception waspikepdf.open(filename) appears as aThe {attr} pikepdf.Page.images property is now deprecated : it only reports images referenced directly by the page and silently omits images drawn thr…
Added {class}pikepdf.JobBuilder, a fluent, Pythonic builder for qpdf jobs.
It assembles a job specification with chained, snake_case methods (input,
output, encrypt, add_pages, split_pages, linearize,
compress, add_attachment, add_overlay, limits, ...) and runs it
via the existing {class}pikepdf.Job, without hand-writing qpdf's camelCase
job JSON. Encryption permissions are expressed with the familiar
{class}pikepdf.Permissions/{class}pikepdf.Encryption models, and a
.set(**kwargs) escape hatch reaches any other job option. Additional
methods cover image optimization (optimize_images,
externalize_inline_images), page/content transforms
(flatten_annotations, flatten_rotation, generate_appearances,
coalesce_contents, normalize_content), content removal
(remove_metadata, remove_info, remove_acroform,
remove_structure, remove_page_labels), page labels
(set_page_labels), version control (min_version, force_version),
and reproducible/inspection helpers (deterministic_id, static_id,
check).
Exposed several pieces of qpdf functionality that pikepdf had not previously
bound:
pikepdf.Pdf.write_qpdf_json,pikepdf.Pdf.from_qpdf_json and {meth}pikepdf.Pdf.update_from_qpdf_jsonqpdf --json-output/--json-input format, version 2). This complements thepikepdf.Object.to_json. Addedpikepdf.JSONStreamData to control how stream data is represented.pikepdf.Pdf.get_xref_table returns the cross-reference table aspikepdf.XrefEntry), complementing the print-onlypikepdf.Pdf.show_xref_table.pikepdf.Pdf.fix_dangling_references repairs references to objectspikepdf.Page.flatten_rotation bakes a page's /Rotate value into itspikepdf.Page.copy_annotations copies annotations (and associated formpikepdf.Page.get_matrix_for_transformations andpikepdf.Page.get_matrix_for_form_xobject_placement expose qpdf'spikepdf.AcroForm.validate,pikepdf.AcroForm.invalidate_cache andpikepdf.AcroForm.transform_annotations for working with interactiveAdded {meth}pikepdf.Page.get_images, which by default recurses into nested
form XObjects to find images. The {attr}pikepdf.Page.images property is now
deprecated: it only reports images referenced directly by the page and
silently omits images drawn through form XObjects, which made it appear as if a
page "has no images" when it clearly did. Use get_images() instead, or
get_images(recursive=False) for the old behavior.
Added {attr}pikepdf.Page.rotation, a property that reports a page's effective
clockwise rotation normalized to [0, 360). Unlike the raw page.Rotate
attribute, it resolves a /Rotate value inherited from the page tree and
reports 0 when no rotation is set, instead of raising. Assigning to it sets
the absolute rotation. This addresses the long-standing confusion between the
page.Rotate attribute and the page.rotate() method (#467).
{meth}pikepdf.Page.rotate now defaults relative to False, so
page.rotate(90) sets an absolute rotation. Passing relative as a
positional argument is deprecated and emits a DeprecationWarning; pass it
as a keyword argument instead, e.g. page.rotate(90, relative=True).
Positional support will be removed in pikepdf 11.
Added {meth}pikepdf.Pdf.add_pages_from to copy pages between documents while
preserving interactive AcroForm form fields, returning a
{class}pikepdf.PageCopyResult. Naive pages.extend() across documents and
save() of documents with orphaned form widgets now emit
{class}pikepdf.PageCopyWarning. (#670, #207)
When copying pages, named destinations referenced by the copied pages'
annotations (e.g. table-of-contents links) are now carried into the
destination document — both the PDF 1.2 Names.Dests name tree and the
legacy PDF 1.1 Root.Dests dictionary — so internal links keep working
regardless of merge order. Name collisions are renamed and reported via
{class}pikepdf.PageCopyResult (named_dests_added, renamed_dests,
dropped_dests). Naive pages.extend() now also warns when copied pages
reference named destinations. (#148)
/Decode array, which caused colors to/Decode such as [1, 0]. {meth}pikepdf.PdfImage.as_pil_image andpikepdf.PdfImage.extract_to now apply /Decode as a linear/Decode was honored only for650apply_decode_array parameter (default True).apply_decode_array=False to retrieve the raw stored sample values/Decode remaps palette indices rather than colors -- a non-identity/Decode there now emits a warning), and DCT (JPEG) / JPX (JPEG 2000)/Decodepikepdf.Pdf.save decompressing streams when called withcompress_streams=False and no explicit stream_decode_level. qpdf 11.10generalized, which caused suchnone in this case, restoring thecompress_streams=False alone does not trigger676.pikepdf.AcroForm.validate binding calls qpdf'sQPDFAcroFormDocumentHelper::validate, which was added in qpdf 12.3.0, soDeleting pages <deleting_pages> topic now explains the behavior and196.Copying metadata between documents <copymetadata> topic, including why188.Added {class} pikepdf.ReferenceCycleError (a subclass of {class} pikepdf.PdfError ), raised when an operation would create a cycle of direct (non-indi
pikepdf.ReferenceCycleError (a subclass ofpikepdf.PdfError), raised when an operation would create a cycle ofpikepdf.Pdf.make_indirect to create apikepdf.sanitize module with curated, low-risk helpers forremove_javascript,remove_attachments, remove_external_access, remove_thumbnails,remove_search_index, remove_multimedia (Rendition/Movie/Sound/RichMedia/3Dremove_web_capture (/SpiderInfo), remove_private_app_dataremove_collection (PDF portfolio view), plus apikepdf.sanitize.Sanitizer builder for chaining these operations. The/GoToE) as external access, and remove_attachments/AF associated-file references from every object (XObjects,sanitize topic673.cp314t) binary wheels toPage's attribute, item and get accessors in C++ instead ofName, Array, Dictionary, the Name.AttrInteger/Boolean/Real, and NamePath)Upgraded to cibuildwheel 3.4.1 and refreshed pinned GitHub Actions ( actions/checkout@v6 , actions/upload-artifact@v7 , actions/download-artifact@v8 ,
actions/checkout@v6, actions/upload-artifact@v7,actions/download-artifact@v8, codecov/codecov-action@v6). Dropped themsvcp140*.dll, vcruntime140*.dll, concrt140.dll)qpdf30.dll. Shipping them caused a second,PyInterpreterState_Get ... the GIL is released (the current Python thread state is NULL) error or anImportError for pikepdf._core, typically only in some launch environmentsdelvewheel vendors avcruntime140, so the wheel remains self-contained without an un-mangled718.pikepdf._core fails to import: it nowDictionary or Array objects that form a cyclic reference graph, for examplea['/Kids'] = [b]; b['/Kids'] = [a]; a == b. The cycle detector previously keyedunparseBinary(), which itself recurses through the whole731.AttributeError when reading a document outline ("bookmarks") whose/Title field. By default, Pdf.open_outline()/Title as an empty string; passingstrict=True raises OutlineStructureError instead. Fixes :issue:730.Fixed a segmentation fault when an object that is not an Encryption , dict , bool , or None (for example a list or unittest.mock.MagicMock ) was passe
Encryption, dict,bool, or None (for example a list or unittest.mock.MagicMock) was passedencryption argument of Pdf.save(). A TypeError is now raised instead.727.Page.add_content_token_filter() if thePdf._token_filter_refs attribute. The attribute is now reset before use.leaked instances/types/functions report atpytest.mark.parametrizePIKEPDF_NANOBIND_LEAK_WARNINGS=1 before importing728.Array.append(None) raising TypeError instead of inserting a PDF725.Dictionary.__setattr__(name, None) (i.e. d.Key = None) raisingTypeError instead of the documented ValueError advising to use del toArray.append(None).Python.h is included before any standard library headers724.Fixed build to continue generating Python version specific wheels for 3.12 and 3.13 due to open issue in nanobind. Fixes :issue: 723 . Thanks @mgorny
723. ThanksNothing published for this version
Released v10.6.0 with version bump only.
Fixed a regression during nanobind migration (exception hierarchy unintentionally changed).
Replaced pybind11 with nanobind and added full freethreading support. pikepdf binary size is now both ~20% smaller and about 10% faster thanks to nano
Nothing published for this version
Fixed logger in ctm module using __file__ instead of __name__, which produced unhelpful log names. 712
ctm module using __file__ instead of __name__,
which produced unhelpful log names. :issue:712Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Manually created due to irrelevant CI build failure
Full Changelog: https://github.com/pikepdf/pikepdf/compare/v9.5.1...v9.5.2
Manually created due to irrelevant CI build failure
Full Changelog: https://github.com/pikepdf/pikepdf/compare/v9.5.0...v9.5.1
Full Changelog: https://github.com/pikepdf/pikepdf/compare/v9.5.0...v9.5.1
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →