NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #319 most downloaded on PyPI
A high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Last release 1 months ago
06 Aug 2026
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
9 years old
138 releases · first in 2017
This fixes #412, #411 and #407 and also inroduces some minor usability enhancements.
This fixes #412, #411 and #407 and also inroduces some minor usability enhancements.
Fixed #412 (“Feature Request: Allow controlling whether TOC entries should be collapsed”)
Fixed #411 (“Seg Fault with page.firstWidget”)
Fixed #407 (“Annot.setOpacity trouble”)
Changed methods Annot.setBorder() , Annot.setColors() , Link.setBorder() , and Link.setColors() to also accept direct parameters, and not just cumbersome dictionaries.
- Added several new methods to the Document class, which make dealing with PDF low-level structures easier. I also decided to provide them as “normal”
Added several new methods to the Document class, which make dealing with PDF low-level structures easier. I also decided to provide them as “normal” methods (as opposed to private ones starting with an underscore “_”). These are Document.xrefObject() , Document.xrefStream() , Document.xrefStreamRaw() , Document.PDFTrailer() , Document.PDFCatalog() , Document.metadataXML() , Document.updateObject() , Document.updateStream() .
Added Tools.mupdf_disply_errors() which sets the display of mupdf errors on sys.stderr .
Added a commandline facility. This a major new feature: you can now invoke several utility functions via “python -m fitz …” . It should obsolete the need for many of the most trivial scripts. Please refer to Module .
One column per quarter.
Minor changes to better synchronize the binary image streams of TextPage image blocks and Document.extractImage() images.
Minor changes to better synchronize the binary image streams of TextPage image blocks and Document.extractImage() images.
Fixed issue #394 (“PyMuPDF Segfaults when using TOOLS.mupdf_warnings()”).
Changed redirection of MuPDF error messages: apart from writing them to Python sys.stderr , they are now also stored with the MuPDF warnings.
Changed Tools.mupdf_warnings() to automatically empty the store (if not deactivated via a parameter).
Changed Page.getImageBbox() to return an infinite rectangle if the image could not be located on the page – instead of raising an exception.
- Fixed issue #390 (“Incomplete deletion of annotations”).
Fixed issue #390 (“Incomplete deletion of annotations”).
Changed Page.searchFor() / Document.searchPageFor() to also support the flags parameter, which controls the data included in a TextPage .
Changed Document.getPageImageList() , Document.getPageFontList() and their Page counterparts to support a new parameter full . If true, the returned items will contain the xref of the Form XObject where the font or image is referenced.
On rare occasions, page rendering and text extraction caused interpreter segmentation faults. When fonts on a page (illegally) were meeting all of the
On rare occasions, page rendering and text extraction caused interpreter segmentation faults. When fonts on a page (illegally) were meeting all of the following conditions:
Identity-H / Identity-V ancodingThen MuPDF reacted with appropriate warning messages (which are intercepted by PyMuPDF). In prior versions, PyMuPDF did not expect non-UTF8 messages and used a too "optimistic" char to unicode conversion method, which brought down the interpreter. This has changed in the new version.
More performance improvements for text extraction.
Fixed second part of issue #381 (see item in v1.16.4).
Added Page.getTextPage() , so it is no longer required to create an intermediate display list for text extractions. Page level wrappers for text extraction and text searching are now based on this, which should improve performance by ca. 5%.
also new generators for document pages and page annotations, widget and links.
also new generators for document pages and page annotations, widget and links.
Fixed issue #381 (“TextPage.extractDICT … failed … after upgrading … to 1.16.3”)
Added method Document.pages() which delivers a generator iterator over a page range.
Added method Page.links() which delivers a generator iterator over the links of a page.
Added method Page.annots() which delivers a generator iterator over the annotations of a page.
Added method Page.widgets() which delivers a generator iterator over the form fields of a page.
Changed Document.is_form_pdf to now contain the number of widgets, and False if not a PDF or this number is zero.
significant performance improvements for dict / rawdict text extraction
dict / rawdict text extractionPage.getText() now support text extraction for "blocks" and "words"Minor changes compared to version 1.16.2. The code of the “dict” and “rawdict” variants of Page.getText() has been ported to C which has greatly improved their performance. This improvement is mostly noticeable with text-oriented documents, where they now should execute almost two times faster.
Fixed issue #369 (“mupdf: cmsCreateTransform failed”) by removing ICC colorspace support.
Changed Page.getText() to accept additional keywords “blocks” and “words”. These will deliver the results of Page.getTextBlocks() and Page.getTextWords() , respectively. So all text extraction methods are now available via a uniform API. Correspondingly, there are now new methods TextPage.extractBLOCKS() and TextPage.extractWords() .
Changed Page.getText() to default bit indicator TEXT_INHIBIT_SPACES to off . Insertion of additional spaces is not suppressed by default.
- Changed text extraction methods of Page to allow detail control of the amount of extracted data.
Changed text extraction methods of Page to allow detail control of the amount of extracted data.
Added planish_line() which maps a given line (defined as a pair of points) to the x-axis.
Fixed an issue (w/o Github number) which brought down the interpreter when encountering certain non-UTF-8 encodable characters while using Page.getText() with the “dict” option.
Fixed issue #362 (“Memory Leak with getText(‘rawDICT’)”).
- Added property Quad.is_convex which checks whether a line is contained in the quad if it connects two points of it.
Added property Quad.is_convex which checks whether a line is contained in the quad if it connects two points of it.
Changed Document.insert_pdf() to now allow dropping or including links and annotations independently during the copy. Fixes issue #352 (“Corrupt PDF data and …”), which seemed to intermittently occur when using the method for some problematic PDF files.
Fixed a bug which, in matrix division using the syntax “m1/m2” , caused matrix “m1” to be replaced by the result instead of delivering a new matrix.
Fixed issue #354 (“SyntaxWarning with Python 3.8”). We now always use “==” for literals (instead of the “is” Python keyword).
Fixed issue #353 (“mupdf version check”), to no longer refuse the import when there are only patch level deviations from MuPDF.
The old names (starting with “ANNOT_*” or “WIDGET_*”) will be available as deprecated synonyms.
This major new version of MuPDF comes with several nice new or changed features. Some of them imply programming API changes, however. This is a synopsis of what has changed:
PDF document encryption and decryption is now fully supported . This includes setting permissions , passwords (user and owner passwords) and the desired encryption method.
In response to the new encryption features, PyMuPDF returns an integer (ie. a combination of bits) for document permissions, and no longer a dictionary.
Redirection of MuPDF errors and warnings is now natively supported. PyMuPDF redirects error messages from MuPDF to sys.stderr and no longer buffers them. Warnings continue to be buffered and will not be displayed. Functions exist to access and reset the warnings buffer.
Annotations are now only supported for PDF .
Annotations and widgets (form fields) are now separate object chains on a page (although widgets technically still are PDF annotations). This means, that you will never encounter widgets when using Page.firstAnnot or Annot.next() . You must use Page.firstWidget and Widget.next() to access form fields.
As part of MuPDF’s changes regarding widgets, only the following four fonts are supported, when adding or changing form fields: Courier, Helvetica, Times-Roman and ZapfDingBats .
List of change details:
Added Document.can_save_incrementally() which checks conditions that are preventing use of option incremental=True of Document.save() .
Added Page.firstWidget which points to the first field on a page.
Added Page.getImageBbox() which returns the rectangle occupied by an image shown on the page.
Added Annot.setName() which lets you change the (icon) name field.
Added outputting the text color in Page.getText() : the “dict” , “rawdict” and “xml” options now also show the color in sRGB format.
Changed Document.permissions to now contain an integer of bool indicators – was a dictionary before.
Changed Document.save() , Document.write() , which now fully support password-based decryption and encryption of PDF files.
Changed the names of all Python constants related to annotations and widgets. Please make sure to consult the Constants and Enumerations chapter if your script is dealing with these two classes. This decision goes back to the dropped support for non-PDF annotations. The old names (starting with “ANNOT_” or “WIDGET_”) will be available as deprecated synonyms.
Changed font support for widgets: only Cour (Courier), Helv (Helvetica, default), TiRo (Times-Roman) and ZaDb (ZapfDingBats) are accepted when adding or changing form fields. Only the plain versions are possible – not their italic or bold variations. Reading widgets, however will show its original font.
Changed the name of the warnings buffer to Tools.mupdf_warnings() and the function to empty this buffer is now called Tools.reset_mupdf_warnings() .
Changed Page.getPixmap() , Document.get_page_pixmap() : a new bool argument annots can now be used to suppress the rendering of annotations on the page.
Changed Page.add_file_annot() and Page.add_text_annot() to enable setting an icon.
Removed widget-related methods and attributes from the Annot object.
Removed Document attributes openErrCode , openErrMsg , and Tools attributes / methods stderr , reset_stderr , stdout , and reset_stdout .
Removed thirdparty zlib dependency in PyMuPDF: there are now compression functions available in MuPDF. Source installers of PyMuPDF may now omit this extra installation step.
No version published for MuPDF v1.15.0
Changes in Version 1.14.20 / 1.14.21
Changed text marker annotations to support multiple rectangles / quadrilaterals. This fixes issue #341 (“Question : How to addhighlight so that a string spread across more than a line is covered by one highlight?”) and similar (#285).
Fixed issue #331 (“Importing PyMuPDF changes warning filtering behaviour globally”).
Nothing published for this version
Nothing published for this version
added method to check PDF signature status
Fixed issue #319 (“InsertText function error when use custom font”).
Added new method Document.get_sigflags() which returns information on whether a PDF is signed. Resolves issue #326 (“How to detect signature in a form pdf?”).
You can now re-layout documents and trace changed page numbers.
You can now re-layout documents and trace changed page numbers.
Nothing published for this version
This is an extension of v1.11.1.
This is an extension of v1.11.1.
New Page.insertFont() creates a PDF /Font object and returns its object number.
New Document.extractFont() extracts the content of an embedded font given its object number.
Methods FontList(…) items no longer contain the PDF generation number. This value never had any significance. Instead, the font file extension is included (e.g. “pfa” for a “PostScript Font for ASCII”), which is more valuable information.
Fonts other than “simple fonts” (Type1) are now also supported.
New options to change Pixmap size:
Method Pixmap.shrink() reduces the pixmap proportionally in place.
A new Pixmap copy constructor allows scaling via setting target width and height.
MuPDF version 1.10 has a significant impact on our bindings. Some of the changes also affect the API – in other words, you as a PyMuPDF user.
MuPDF v1.10 Impact
MuPDF version 1.10 has a significant impact on our bindings. Some of the changes also affect the API – in other words, you as a PyMuPDF user.
Link destination information has been reduced. Several properties of the linkDest class no longer contain valuable information. In fact, this class as a whole has been deleted from MuPDF’s library and we in PyMuPDF only maintain it to provide compatibility to existing code.
In an effort to minimize memory requirements, several improvements have been built into MuPDF v1.10:
A new config.h file can be used to de-select unwanted features in the C base code. Using this feature we have been able to reduce the size of our binary _fitz.o / _fitz.pyd by about 50% (from 9 MB to 4.5 MB). When UPX-ing this, the size goes even further down to a very handy 2.3 MB.
The alpha (transparency) channel for pixmaps is now optional. Letting alpha default to False significantly reduces pixmap sizes (by 20% – CMYK, 25% – RGB, 50% – GRAY). Many Pixmap constructors therefore now accept an alpha boolean to control inclusion of this channel. Other pixmap constructors (e.g. those for file and image input) create pixmaps with no alpha altogether. On the downside, save methods for pixmaps no longer accept a savealpha option: this channel will always be saved when present. To minimize code breaks, we have left this parameter in the call patterns – it will just be ignored.
DisplayList and TextPage class constructors now require the mediabox of the page they are referring to (i.e. the page.bound() rectangle). There is no way to construct this information from other sources, therefore a source code change cannot be avoided in these cases. We assume however, that not many users are actually employing these rather low level classes explicitly. So the impact of that change should be minor.
Other Changes compared to Version 1.9.3
The new Document method write() writes an opened PDF to memory (as opposed to a file, like save() does).
An annotation can now be scaled and moved around on its page. This is done by modifying its rectangle.
Annotations can now be deleted. Page contains the new method deleteAnnot() .
Various annotation attributes can now be modified, e.g. content, dates, title (= author), border, colors.
Method Document.insert_pdf() now also copies annotations of source pages.
The Pages class has been deleted. As documents can now be accessed with page numbers as indices (like doc[n] = doc.loadPage(n) ), and document object can be used as iterators, the benefit of this class was too low to maintain it. See the following comments.
loadPage(n) / doc[n] now accept arbitrary integers to specify a page number, as long as n < pageCount . So, e.g. doc[-500] is always valid and will load page (-500) % pageCount .
A document can now also be used as an iterator like this: for page in doc: …<do something with “page”> … . This will yield all pages of doc as page .
The Pixmap method getSize() has been replaced with property size . As before Pixmap.size == len(Pixmap) is true.
In response to transparency (alpha) being optional, several new parameters and properties have been added to Pixmap and Colorspace classes to support determining their characteristics.
The Page class now contains new properties firstAnnot and firstLink to provide starting points to the respective class chains, where firstLink is just a mnemonic synonym to method loadLinks() which continues to exist. Similarly, the new property rect is a synonym for method bound() , which also continues to exist.
Pixmap methods samplesRGB() and samplesAlpha() have been deleted because pixmaps can now be created without transparency.
Rect now has a property irect which is a synonym of method round() . Likewise, IRect now has property rect to deliver a Rect which has the same coordinates as floats values.
Document has the new method searchPageFor() to search for a text string. It works exactly like the corresponding Page.searchFor() with page number as additional parameter.
Version 1.9.2 is based on MuPDF 1.9a as was version 1.9.1. Its major new features are around PDF creation - most prominently Document.insertPDF().
Version 1.9.2 is based on MuPDF 1.9a as was version 1.9.1.
Its major new features are around PDF creation - most prominently Document.insertPDF().
It also contains minor bug fixes and improvements.
This version is also based on MuPDF v1.9a. Changes compared to version 1.9.1:
fitz.open() (no parameters) creates a new empty PDF document, i.e. if saved afterwards, it must be given a .pdf extension.
Document now accepts all of the following formats ( Document and open are synonyms):
open() ,
open(filename) (equivalent to open(filename, None) ),
open(filetype, area) (equivalent to open(filetype, stream = area) ).
Type of memory area stream may be bytes or bytearray . Thus, e.g. area = open(“file.pdf”, “rb”).read() may be used directly (without first converting it to bytearray).
New method Document.insert_pdf() (PDFs only) inserts a range of pages from another PDF.
Document objects doc now support the len() function: len(doc) == doc.pageCount .
New method Document.getPageImageList() creates a list of images used on a page.
New method Document.getPageFontList() creates a list of fonts referenced by a page.
New pixmap constructor fitz.Pixmap(doc, xref) creates a pixmap based on an opened PDF document and an xref number of the image.
New pixmap constructor fitz.Pixmap(cspace, spix) creates a pixmap as a copy of another one spix with the colorspace converted to cspace . This works for all colorspace combinations.
Pixmap constructor fitz.Pixmap(colorspace, width, height, samples) now allows samples to also be bytes , not only bytearray .
Your coding agent can read these notes before it upgrades. Set up the MCP server →