NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #296 most downloaded on PyPI
A high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Last release 1 months ago
06 Aug 2026
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
9 years old
138 releases · first in 2017
PyMuPDF-1.23.2 has been released.
PyMuPDF-1.23.2 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.23.2 (2023-08-28)
Nothing published for this version
One column per quarter.
PyMuPDF-1.23.1 has been released.
PyMuPDF-1.23.1 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.23.1 (2023-08-24)
Updated README and package summary description.
Fixed a problem on some Linux installations with Python-3.10
(and possibly earlier versions) where import fitz failed with
ImportError: libcrypt.so.2: cannot open shared object file: No such file or directory.
Fixed incompatible architecture error on MacOS arm64.
Fixed installation warning from Poetry about missing entry in wheels' RECORD files.
PyMuPDF-1.23.0 has been released.
PyMuPDF-1.23.0 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.23.0 (2023-08-22)
Add method find_tables() to the Page object.
This allows locating tables on any supported document page, and extracting table content by cell.
New "rebased" implementation of PyMuPDF.
The rebased implementation is available as Python module
fitz_new. It can be used as a drop-in replacement with import fitz_new as fitz.
Python-independent MuPDF libraries are now in a second wheel called
PyMuPDFb that will be automatically installed by pip.
This is to save space on pypi.org - a full release only needs one
PyMuPDFb wheel for each OS.
Bug fixes:
Other changes:
Dropped support for Python-3.7.
Fix for wrong page / annot /Contents cleaning.
We need to set pdf_filter_options::no_update to zero.
Added new function get_tessdata().
Cope with problem /Annot arrays.
When copying page annotations in method Document.insert_pdf we
previously did not check the validity of members of the /Annots
array. For faulty members (like null or non-dictionary items) this
could cause unnecessary exceptions. This fix implements more checks
and skips such array items.
Additional annotation type checks.
We did not previously check for annotation type when getting / setting annotation border properties. This is now checked in accordance with MuPDF.
Increase fault tolerance.
Avoid exceptions in method insert_pdf() when source pages contains
invalid items in the /Annots array.
Return empty border dict for applicable annots.
We previously were returning a non-empty border dictionary even for
non-applicable annotation types. We now return the empty dictionary
{} in these cases. This requires some corresponding changes in the
annotation .update() method, namely for dashes and border width.
Restrict set_rect to applicable annot types.
We were insufficiently excluding non-applicable annotation types
from set_rect() method. We now let MuPDF catch unsupported
annotations and return False in these cases.
Wrong fontsize computation in page.get_texttrace().
When computing the font size we were using the final text
transformation matrix, where we should have taken span->trm
instead. This is corrected here.
Updates to cope with changes to latest MuPDF.
pdf_lookup_anchor() has been removed.
Update fill_textbox to better respect rect.width
The function norm_words in fill_textbox had a bug in its last loop, appending n+1 characters when actually measuring width of n characters. It led to a bug in fill_texbox when you tried to write a single word mostly composed of "wide" letters (M,m, W, w...), causing the written text to exceed the given rect.
The fix was just to replace n+1 by n.
Add script_focus and script_blur options to widget.
Nothing published for this version
Nothing published for this version
PyMuPDF-1.22.5 has been released.
PyMuPDF-1.22.5 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
` python -m pip install --upgrade pymupdf `
Changes in version 1.22.5 (2023-06-21)
This release uses MuPDF-1.22.2.
Bug fixes:
Fixed #2365
Fixed #2391
Fixed #2400
Fixed #2404
Fixed #2430
Fixed #2450
Fixed #2462
Fixed #2468
New features:
Changed Annotations now support "cloudy" borders. The Annot.border property has the new item clouds, and method Annot.set_border supports the corresponding clouds argument.
Changed Radio button widgets in the same RB group are now consistently updated if the group is defined in the standard way.
Added Support for the /Locked key in PDF Optional Content. This array inside the catalog entry /OCProperties can now be extracted and set.
Added Support for new parameter tessdata in OCR functions. New function get_tessdata locates the language support folder if Tesseract is installed.
This release uses MuPDF-1.22.2 .
Bug fixes:
Fixed #2365 : Incorrect dictionary values for type “fs” drawings.
Fixed #2391 : Check box automatically uncheck when we update same checkbox more than 1 times.
Fixed #2400 : Gaps within text of same line not filled with spaces.
Fixed #2404 : Blacklining an image in PDF won’t remove underlying content in version 1.22.X.
Fixed #2430 : Incorrectly reducing ref count of Py_None.
Fixed #2450 : Empty fill color and fill opacity for paths with fill and stroke operations with 1.22.*
Fixed #2462 : Error at “get_drawing(extended=True )”
Fixed #2468 : Decode error when trying to get drawings
Fixed #2710 : page.rect and text location wrong / differing from older version
Fixed #2723 : When will a Python 3.12 wheel be available?
New features:
Changed Annotations now support “cloudy” borders. The Annot.border property has the new item clouds , and method Annot.set_border() supports the corresponding clouds argument.
Changed Radio button widgets in the same RB group are now consistently updated if the group is defined in the standard way .
Added Support for the /Locked key in PDF Optional Content. This array inside the catalog entry /OCProperties can now be extracted and set.
Added Support for new parameter tessdata in OCR functions. New function get_tessdata() locates the language support folder if Tesseract is installed.
PyMuPDF-1.22.3 has been released.
PyMuPDF-1.22.3 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.22.3 (2023-05-10)
Bug fixes:
This release uses MuPDF-1.22.0 .
Bug fixes:
Fixed #2333 : Unable to set any of button radio group in form
PyMuPDF-1.22.2 has been released.
PyMuPDF-1.22.2 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.22.2 (2023-04-26)
This release uses MuPDF-1.22.0.
Bug fixes:
PyMuPDF-1.22.1 has been released.
PyMuPDF-1.22.1 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.22.1 (2023-04-18)
This release uses MuPDF-1.22.0.
Bug fixes:
PyMuPDF-1.22.0 has been released.
PyMuPDF-1.22.0 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in version 1.22.0 (2023-04-14)
This release uses MuPDF-1.22.0.
Behavioural changes:
Bug fixes:
page.add_highlight_annot(start=pointa, stop=pointb) not workingColorspace used in Pixmappixmap.tint_with()apply_redactions() move pdf text to right after redactionPage.delete_image() | object has no attribute is_imageStory.element_positions() if callback function prototype is wrongAnnot.get_text("words") - doesn't return the first line of wordsOther:
Add key "/AS (Yes)" to the underlying annot object of a selected button form field.
Remove unused Document methods has_xref_streams() and
has_old_style_xrefs() as MuPDF equivalents have been removed.
Add new Document methods and properties for getting/setting
/PageMode, /PageLayout and /MarkInfo.
New Document property version_count, which contains the number of
incremental saves plus one.
New Document property is_fast_webaccess which tells whether the
document is linearized.
DocumentWriter is now a context manager.
Add support for Pixmap JPEG output.
Add support for drawing rectangles with rounded corners.
get_drawings(): added optional extended arg.
Fixed issue where trace devices' state was not being initialised
correctly; data returned from things like fitz.Page.get_texttrace()
might be slightly altered, e.g. linewidth values.
Output warning to stderr if it looks like we are being used with
current directory containing an invalid fitz/ directory, because
this can break import of fitz module. For example this happens
if one attempts to use fitz when current directory is a PyMuPDF
checkout.
Documentation:
General rework:
Improve insert_file() documentation.
get_bboxlog(): aded optional layers to get_bboxlog().
Page.get_texttrace(): add new dictionary key layer, name of Optional Content Group.
Mention use of Python venv in installation documentation.
Added missing fix for #2057 to release 1.21.1's changelog.
Fixes many links to the PyMuPDF-Utilities repo scripts.
Avoid duplication of changes.txt and docs/changes.rst.
Build
pyproject.toml file to improve builds using pip etc.PyMuPDF-1.21.1 has been released.
PyMuPDF-1.21.1 has been released.
Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
python -m pip install --upgrade pymupdf
Changes in Version 1.21.1 (2022-12-13)
This release uses MuPDF-1.21.1.
Bug fixes:
Other
Page.replace_image() and Page.delete_image().Documentation:
tests/README.md.docs/installation.rst, mention incompatibility with chocolatey.org on Windows.Annot.file_info.PyMuPDF-1.21.0 has been released.
PyMuPDF-1.21.0 has been released.
This release uses MuPDF-1.21.0.
New feature: Stories.
Added wheels for Python-3.11.
Bug fixes:
Document.delete_pages() declines keyword arguments.page.apply_redactions().fontname="Helvetica" can silently fail.draw_rect(): does not respect width if color is not specified.subset_fonts(): make it possible to silence the stdout.clean=True.pdfocr_save() Hard Crash.get_drawings().block_no and block_type switched in get_text() docs.Misc changes to core code:
Other:
swig-python on MacOS for #875.Moved old docs/faq.rst into separate docs/recipes-* files.
docs/faq.rst into separate docs/recipes-* files.docs/installation.rst.JM_MEMORY in docs/tools.rst.docs/app3.rst.Wheels for Windows, Linux and MacOS, and the sdist, are available on pypi.org and can be installed in the usual way, for example:
pip install --upgrade pymupdf
This release uses MuPDF-1.20.3 .
Fixed #1787 . Fix linking issues on Unix systems.
Fixed #1824 . SegFault when applying redactions overlapping a transparent image. (Fixed in MuPDF-1.20.3 .)
Improvements to documentation:
Improved information about building from source in docs/installation.rst .
Clarified memory allocation setting JM_MEMORY` in ``docs/tools.rst .
Fixed link to PDF Reference manual in docs/app3.rst .
Fixed building of html documentation on OpenBSD.
Moved old docs/faq.rst into separate docs/recipes-* files.
Removed some unused files and directories:
installation/
docs/wheelnames.txt
Fix https://github.com/pymupdf/PyMuPDF/pull/1724.
Fixed #1724 . Fix for building on FreeBSD.
Fixed #1771 . linkDest() had a broken call to re.match() , introduced in 1.20.0.
Fixed #1751 . get_drawings() and get_cdrawings() previously always returned with closePath=False .
Fixed #1645 . Default FreeText annotation text color is now black.
Improvements to sphinx-generated documentation:
Use readthedocs theme with enhancements.
Renamed the .txt files to have .rst suffixes.
This release integrates the recently-released MuPDF-1.20.0, and has fixes for #1733 and #1738. The latter also contains an additional fix for occasion
This release integrates the recently-released MuPDF-1.20.0, and has fixes for #1733 and #1738. The latter also contains an additional fix for occasional SEGVs when freeing documents.
Building from source works slightly differently from before:
setup.py for details.This release uses MuPDF-1.20.0 , released 2022-06-15.
Cope with new MuPDF link uri format, changed from #<int>,<int>,<int> to #page=<int>&zoom=<float>,<float>,<float> .
In tests/test_insertpdf.py , use new reference output joined-1.20.pdf . We also check that new output values are approximately the same as the old ones.
Fixed #1738 . Leak of pdf_graft_map . Also fixed a SEGV issue that this seemed to expose, caused by incorrect freeing of underlying fz_document.
Fixed #1733 . Fixed ownership of Annotation.get_pixmap() .
Changes to build/release process:
If pip builds from source because an appropriate wheel is not available, we no longer require MuPDF to be pre-installed. Instead the required MuPDF source is embedded in the sdist and automatically built into PyMuPDF.
Various changes to setup.py to download the required MuPDF release as required. See comments at start of setup.py for details.
Added .github/workflows/build_wheels.yml to control building of wheels on Github.
new method Page.load_widget() to load a widget from its xref
Fixes: #1620, #1601
Enhancements:
Page.load_widget() to load a widget from its xrefpdfcolor which contains 500 predefined PDF colorsQuad class supports operator algebraPage.annots() and Page.widgets() now prohibit reloading the page within their scopeTools class and redefined them as standalonenew in Document.update_stream() is now obsolete.Fixed #1620 . The TextPage created by Page.get_textpage() will now be freed correctly (removed memory leak).
Fixed #1601 . Document open errors should now be more concise and easier to interpret. In the course of this, two PyMuPDF-specific Python exceptions have been added:
EmptyFileError – raised when trying to create a Document ( fitz.open() ) from an empty file or zero-length memory.
FileDataError – raised when MuPDF encounters irrecoverable document structure issues.
Added Page.load_widget() given a PDF field’s xref.
Added Dictionary pdfcolor which provide the about 500 colors defined as PDF color values with the lower case color name as key.
Added algebra functionality to the Quad class. These objects can now also be added and subtracted among themselves, and be multiplied by numbers and matrices.
Added new constants defining the default text extraction flags for more comfortable handling. Their naming convention is like TEXTFLAGS_WORDS for page.get_text("words") . See Text Extraction Flags Defaults .
Changed Page.annots() and Page.widgets() to detect and prevent reloading the page (illegally) inside the iterator loops via Document.reload_page() . Doing this brings down the interpreter. Documented clean ways to do annotation and widget mass updates within properly designed loops.
Changed several internal utility functions to become standalone (“SWIG inline”) as opposed to be part of the Tools class. This, among other things, increases the performance of geometry object creation.
Changed Document.update_stream() to always accept stream updates - whether or not the dictionary object behind the xref already is a stream. Thus the former new parameter is now ignored and will be removed in v1.20.0.
Fixes: #1583, #1552, #1550, #1521, #1518, #1513, #1510, #1417, #1550. Also fixed some undocumented errors that caused the span["origin"] to be incorre
Fixes: #1583, #1552, #1550, #1521, #1518, #1513, #1510, #1417, #1550.
Also fixed some undocumented errors that caused the span["origin"] to be incorrectly set in corner cases.
Added new items "orientation" and associated transformtion matrix to the output of fitz.image_properties(), which contains EXIF data of supporting image files.
A new method Document.xref_copy() allows making xref objects duplicates of each other.
Fixed #1518 . A limited “fix”: in some cases, rectangles and quadrupels were not correctly encoded to support re-drawing by Shape .
Fixed #1521 . This had the same ultimate reason behind issue #1510.
Fixed #1513 . Some Optional Content functions did not support non-ASCII characters.
Fixed #1510 . Support more soft-mask image subtypes.
Fixed #1507 . Immunize against items in the outlines chain, that are "null" objects.
Fixed re-opened #1417 . (“too many open files”). This was due to insufficient calls to MuPDF’s fz_drop_document() . This also fixes #1550 .
Fixed several undocumented issues in relation to incorrectly setting the text span origin point_like .
Fixed undocumented error computing the character bbox in method Page.get_texttrace() when text is flipped (as opposed to just rotated).
Added items to the dictionary returned by image_properties() : orientation and transform report the natural image orientation (EXIF data).
Added method Document.xref_copy() . It will make a given target PDF object an exact copy of a source object.
Fixes: #1505, #1484, #1479, #1474.
Fixes: #1505, #1484, #1479, #1474.
Changes:
/ArtBox etc.Document.xref_set_key() such that dictionary keys will physically be removed if set to value "null".Document.extract_font() to optionally return a dictionary (instead of a tuple).Fixed #1505 . Immunize against circular outline items.
Fixed #1484 . Correct CropBox coordinates are now returned in all situations.
Fixed #1479 .
Fixed #1474 . TextPage objects are now properly deleted again.
Added Page methods and attributes for PDF /ArtBox , /BleedBox , /TrimBox .
Added global attribute TESSDATA_PREFIX for easy checking of OCR support.
Changed Document.xref_set_key() such that dictionary keys will physically be removed if set to value "null" .
Changed Document.extract_font() to optionally return a dictionary (instead of a tuple).
New or changed Pixmap methods color_topusage(), color_count(), warp(). Some of them solve #1397.
Fixes: #1351, #1417, #1418, #1430, #1433
color_topusage(), color_count(), warp(). Some of them solve #1397.irt_xref, set_irt_xref(). Implements #1450.Rect / IRect method torect() which creates a matrix to transform between given rectangles.Page.get_texttrace() now also supports non-horizontal text.This patch version implements minor improvements for Pixmap and also some important fixes.
Fixed #1351 . Reverted code that introduced the memory growth in v1.18.15.
Fixed #1417 . Developed circumvention for growth of open file handles using Document.insert_pdf() .
Fixed #1418 . Developed circumvention for memory growth using Document.insert_pdf() .
Fixed #1430 . Developed circumvention for mass pixmap generations of document pages.
Fixed #1433 . Solves a bbox error for some Type 3 font in PyMuPDF text processing.
Added Pixmap.color_topusage() to determine the share of the most frequently used color. Solves #1397 .
Added Pixmap.warp() which makes a new pixmap from a given arbitrary convex quad inside the pixmap.
Added Annot.irt_xref and Annot.set_irt_xref() to inquire or set the /IRT (“In Response To”) property of an annotation. Implements #1450 .
Added Rect.torect() and IRect.torect() which compute a matrix that transforms to a given other rectangle.
Changed Pixmap.color_count() to also return the count of each color.
Changed Page.get_texttrace() to also return correct span and character bboxes if span["dir"] != (1, 0) .
Page.get_drawings() now includes area orientation for rectangles
Improvements:
Page.get_drawings() now includes area orientation for rectangles"dpi"Fixes: #1388, #1375, #1364, #1342, #1355, #1397, #1408.
This patch version implements minor improvements for Page.get_drawings() and also some important fixes.
Fixed #1388 . Fixed intermittent memory corruption when insert or updating annotations.
Fixed #1375 . Inconsistencies between line numbers as returned by the “words” and the “dict” options of Page.get_text() have been corrected.
Fixed #1364 . The check for being a "rawdict" span in recover_span_quad() now works correctly.
Fixed #1342 . Corrected the check for rectangle infiniteness in Page.show_pdf_page() .
Changed Page.get_drawings() , Page.get_cdrawings() to return an indicator on the area orientation covered by a rectangle. This implements #1355 . Also, the recognition rate for rectangles and quads has been significantly improved.
Changed all text search and extraction methods to set the new flags option TEXT_MEDIABOX_CLIP to ON by default. That bit causes the automatic suppression of all characters that are completely outside a page’s mediabox (in as far as that notion is supported for a document type). This eliminates the need for using clip=page.rect or similar for omitting text outside the visible area.
Added parameter "dpi" to Page.get_pixmap() and Annot.get_pixmap() . When given, parameter "matrix" is ignored, and a Pixmap with the desired dots per inch is created.
Added attributes Pixmap.is_monochrome and Pixmap.is_unicolor allowing fast checks of pixmap properties. Addresses #1397 .
Added method Pixmap.color_count() to determine the unique colors in the pixmap.
Added boolean parameter "compress" to PDF document method Document.update_stream() . Addresses / enables solution for #1408 .
OCR of a document page has been improved a lot compared to v1.19.0. Text extractions now also come with an integrated sort. Fixes: #1328
OCR of a document page has been improved a lot compared to v1.19.0. Text extractions now also come with an integrated sort. Fixes: #1328
This is the first patch version to support MuPDF v1.19.0. Apart from one bug fix, it includes important improvements for OCR support and the option to sort extracted text to the standard reading order “from top-left to bottom-right”.
Fixed #1328 . “words” text extraction again returns correct (x0, y0) coordinates.
Changed Page.get_textpage_ocr() : it now supports parameter dpi to control OCR quality. It is also possible to choose whether the full page should be OCRed or only the images displayed by the page.
Changed Page.get_drawings() and Page.get_cdrawings() to automatically convert colors to RGB color tuples. Implements #1332 . Similar change was applied to Page.get_texttrace() .
Changed Page.get_text() to support a parameter sort . If set to True the output is conveniently sorted.
Introduces major new features like PDF journalling and OCR support by directly invoking Tesseract-OCR. In addition, it is possible to detect whether o
Introduces major new features like PDF journalling and OCR support by directly invoking Tesseract-OCR. In addition, it is possible to detect whether object are covered (hidden) by other objects.
As part of the new version, the following issues have resolved: #1313, #1311, #1290, #1286, #1287, #1284.
This is the first version supporting MuPDF 1.19., published 2021-10-05. It introduces many new features compared to the previous version 1.18..
PyMuPDF has now picked up integrated Tesseract OCR support, which was already present in MuPDF v1.18.0.
Supported images can be OCRed via their Pixmap which results in a 1-page PDF with a text layer.
All supported document pages (i.e. not only PDFs), can be OCRed using specialized text extraction methods. The result is a mixture of standard and OCR text (depending on which part of the page was deemed to require OCRing) that can be searched and extracted without restrictions.
All this requires an independent installation of Tesseract. MuPDF actually (only) needs the location of Tesseract’s "tessdata" folder, where its language support data are stored. This location must be available as environment variable TESSDATA_PREFIX .
A new MuPDF feature is journalling PDF updates , which is also supported by this PyMuPDF version. Changes may be logged, rolled back or replayed, allowing to implement a whole new level of control over PDF document integrity – similar to functions present in modern database systems.
A third feature (unrelated to the new MuPDF version) includes the ability to detect when page objects cover or hide each other . It is now e.g. possible to see that text is covered by a drawing or an image.
Changed terminology and meaning of important geometry concepts: Rectangles are now characterized as finite , valid or empty , while the definitions of these terms have also changed. Rectangles specifically are now thought of being “open”: not all corners and sides are considered part of the rectangle. Please do read the Rect section for details.
Added new parameter "no_new_id" to Document.save() / Document.tobytes() methods. Use it to suppress updating the second item of the document /ID which in PDF indicates that the original file has been updated. If the PDF has no /ID at all yet, then no new one will be created either.
Added a journalling facility for PDF updates. This allows logging changes, undoing or redoing them, or saving the journal for later use. Refer to Document.journal_enable() and friends.
Added new Pixmap methods Pixmap.pdfocr_save() and Pixmap.pdfocr_tobytes() , which generate a 1-page PDF containing the pixmap as PNG image with OCR text layer.
Added Page.get_textpage_ocr() which executes optical character recognition for the page, then extracts the results and stores them together with “normal” page content in a TextPage . Use or reuse this object in subsequent text extractions and text searches to avoid multiple efforts. The existing text search and text extraction methods have been extended to support a separately created textpage – see next item.
Added a new parameter textpage to text extraction and text search methods. This allows reuse of a previously created TextPage and thus achieves significant runtime benefits – which is especially important for the new OCR features. But “normal” text extractions can definitely also benefit.
Added Page.get_texttrace() , a technical method delivering low-level text character properties. It was present before as a private method, but the author felt it now is mature enough to be officially available. It specifically includes a “sequence number” which indicates the page appearance build operation that painted the text.
Added Page.get_bboxlog() which delivers the list of rectangles of page objects like text, images or drawings. Its significance lies in its sequence: rectangles intersecting areas with a lower index are covering or hiding them.
Changed methods Page.get_drawings() and Page.get_cdrawings() to include a “sequence number” indicating the page appearance build operation that created the drawing.
Fixed #1311 . Field values in comboboxes should now be handled correctly.
Fixed #1290 . Error was caused by incorrect rectangle emptiness check, which is fixed due to new geometry logic of this version.
Fixed #1286 . Text alignment for redact annotations is working again.
Fixed #1287 . Infinite loop issue for non-Windows systems when applying some redactions has been resolved.
Fixed #1284 . Text layout destruction after applying redactions in some cases has been resolved.
Changes in Version 1.18.18 / 1.18.19
Fixed issue #1266 . Failure to set Pixmap.samples in important cases, was hotfixed in a new version 1.18.19.
Fixed issue #1257 . Removing the read-only flag from PDF fields is now possible.
Fixed issue #1252 . Now correctly specifying the zoom value for PDF link annotations.
Fixed issue #1244 . Now correctly computing the transform matrix in Page.get_image__bbox() .
Fixed issue #1241 . Prevent returning artifact characters in Page.get_textbox() , which happened in certain constellations.
Fixed issue #1234 . Avoid creating infinite rectangles in corner cases – Page.get_drawings() , Page.get_cdrawings() .
Added test data and test scripts to the source PyPI source distribution.
Nothing published for this version
This version fixes #1257, #1252, #1244, #1241, #1234, #1236, #1227.
This version fixes #1257, #1252, #1244, #1241, #1234, #1236, #1227.
Focus of this version are major performance improvements of selected functions.
Focus of this version are major performance improvements of selected functions.
Fixed issue #1199 . Using a non-existing page number in Document.get_page_images() and friends will no longer lead to segfaults.
Changed Page.get_drawings() to now differentiate between “stroke”, “fill” and combined paths. Paths containing more than one rectangle (i.e. “re” items) are now supported. Extracting “clipped” paths is now available as an option.
Added Page.get_cdrawings() , performance-optimized version of Page.get_drawings() .
Added Pixmap.samples_mv , memoryview of a pixmap’s pixel area. Does not copy and thus always accesses the current state of that area.
Added Pixmap.samples_ptr , Python “pointer” to a pixmap’s pixel area. Allows much faster creation (factor 800+) of Qt images.
The fitz module now supports text extraction via a new subcommand "gettext". Among a couple of modes, preserving the original layout can be chosen.
The fitz module now supports text extraction via a new subcommand "gettext". Among a couple of modes, preserving the original layout can be chosen.
Also fixed #1187, #1184, #1154, #1152 and #1146.
Fixed issue #1184 . Existing PDF widget fonts in a PDF are now accepted (i.e. not forcedly changed to a Base-14 font).
Fixed issue #1154 . Text search hits should now be correct when clip is specified.
Fixed issue #1152 .
Fixed issue #1146 .
Added Link.flags and Link.set_flags() to the Link class. Implements enhancement requests #1187 .
Added option to simulate TextWriter.fill_textbox() output for predicting the number of lines, that a given text would occupy in the textbox.
Added text output support as subcommand gettext to the fitz CLI module. Most importantly, original physical text layout reproduction is now supported.
Apart from some minor fixes, this release introduces support for small caps in TextWriter based text output.
Apart from some minor fixes, this release introduces support for small caps in TextWriter based text output.
In addition, method Document.subset_fonts() now prefixes subsetted font names with the 6 upper case letter prefix as prescribed by the PDF standard.
List of fixed issues: #1088, #1081, #1078, #1085.
Fixed issue #1088 . Removing an annotation’s fill color should now work again both ways, using the fill_color=[] argument in Annot.update() as well as fill=[] in Annot.set_colors() .
Fixed issue #1081 . Document.subset_fonts() : fixed an error which created wrong character widths for some fonts.
Fixed issue #1078 . Page.get_text() and other methods related to text extraction: changed the default value of the TextPage flags parameter. All whitespace and ligatures are now preserved.
Fixed issue #1085 . The old snake_cased alias of fitz.detTextlength is now defined correctly.
Changed Document.subset_fonts() will now correctly prefix font subsets with an appropriate six letter uppercase tag, complying with the PDF specification.
Added new method Widget.button_states() which returns the possible values that a button-type field can have when being set to “on” or “off”.
Added support of text with Small Capital letters to the Font and TextWriter classes. This is reflected by an additional bool parameter small_caps in various of their methods.
undocumented occasional error calculating envelopping rectangle for paths in Page.get_drawings()
The following habe been fixed:
Page.get_drawings()TextWriter.fill_textbox()Font.char_lengths() which returns a tuple of all character widths for a given string. An improved version of Font.text_length()Font.text_length()del statementDocument.del_toc_item(): the item's title text will no longer be removed - instead the item is shown grayed-out to indicate its deletion.Finished implementing new, “snake_cased” names for methods and properties, that were “camelCased” and awkward in many aspects. At the end of this documentation, there is section Deprecated Names with more background and a mapping of old to new names.
Fixed issue #1053 . Page.insert_image() : when given, include image mask in the hash computation.
Fixed issue #1043 . Added Pixmap.getPNGdata to the aliases of Pixmap.tobytes() .
Fixed an internal error when computing the enveloping rectangle of drawn paths as returned by Page.get_drawings() .
Fixed an internal error occasionally causing loops when outputting text via TextWriter.fill_textbox() .
Added Font.char_lengths() , which returns a tuple of character widths of a string.
Added more ways to specify pages in Document.delete_pages() . Now a sequence (list, tuple or range) can be specified, and the Python del statement can be used. In the latter case, Python slices are also accepted.
Changed Document.del_toc_item() , which disables a single item of the TOC: previously, the title text was removed. Instead, now the complete item will be shown grayed-out by supporting viewers.
Method Page.insert_image has been rewritten for improved performance in standard cases. Also introduced option to re-use pre-existing images in the fi
Method Page.insert_image has been rewritten for improved performance in standard cases. Also introduced option to re-use pre-existing images in the file directly to provide another performance boost.
Other changes:
Fixed issue #1014 .
Fixed an internal memory leak when computing image bboxes – Page.get_image_bbox() .
Added support for low-level access and modification of the PDF trailer. Applies to Document.xref_get_keys() , Document.xref_get_key() , and Document.xref_set_key() .
Added documentation for maintaining private entries in PDF metadata.
Added documentation for handling transparent image insertions, Page.insert_image() .
Added Page.get_image_rects() , an improved version of Page.get_image_bbox() .
Changed Document.delete_pages() to support various ways of specifying pages to delete. Implements #1042 .
Changed Page.insert_image() to also accept the xref of an existing image in the file. This allows “copying” images between pages, and extremely fast multiple insertions.
Changed Page.insert_image() to also accept the integer parameter alpha . To be used for performance improvements.
Changed Pixmap.set_alpha() to support new parameters for pre-multiplying colors with their alpha values and setting a specific color to fully transparent (e.g. white).
Changed Document.embfile_add() to automatically set creation and modification date-time. Correspondingly, Document.embfile_upd() automatically maintains modification date-time ( /ModDate PDF key), and Document.embfile_info() correspondingly reports these data. In addition, the embedded file’s associated “collection item” is included via its xref . This supports the development of PDF portfolio applications.
Changes in Version 1.18.11 / 1.18.12
Fixed issue #972 . Improved layout of source distribution material.
Fixed issue #962 . Stabilized Linux distribution detection for generating PyMuPDF from sources.
Added: Page.get_xobjects() delivers the result of Document.get_page_xobjects() .
Added: Page.get_image_info() delivers meta information for all images shown on the page.
Added: Tools.mupdf_display_warnings() allows setting on / off the display of MuPDF-generated warnings. The default is off.
Added: Document.ez_save() convenience alias of Document.save() with some different defaults.
Changed: Image extractions of document pages now also contain the image’s transformation matrix . This concerns Page.get_image_bbox() and the DICT, JSON, RAWDICT, and RAWJSON variants of Page.get_text() .
Nothing published for this version
Meta information for images embedded in document pages has been enriched by the so-called transformation matrix. It can be used to find out, what "hap
Meta information for images embedded in document pages has been enriched by the so-called transformation matrix. It can be used to find out, what "happened" to the image rectangle to make it fit in its bbox on the page, like scaling and rotation.
Other changes are mostly minor bug fixes: #990 #972
A new Page method get_image_info() is also available, which extracts image meta information from the page's TextPage - much like the corresponding Page.get_text("dict"), but without extracting any text or the image binary data themselves.
included PDF trailer access in Document.xref_get_key()
Fixed: #941 #929 #927
Document.xref_get_key()Fixed issue #941 . Added old aliases for DisplayList.get_pixmap() and DisplayList.get_textpage() .
Fixed issue #929 . Stabilized removal of JavaScript objects with Document.scrub() .
Fixed issue #927 . Removed a loop in the reworked TextWriter.fill_textbox() .
Changed Document.xref_get_keys() and Document.xref_get_key() to also allow accessing the PDF trailer dictionary. This can be done by using -1 as the xref number argument.
Added a number of functions for reconstructing the quads for text lines, spans and characters extracted by Page.get_text() options “dict” and “rawdict”. See recover_quad() and friends.
Added Tools.unset_quad_corrections() to suppress character quad corrections (occasionally required for erroneous fonts).
Fixed #888, #895, #896, #885, #922 Implemented #897 (text output right-to-left).
Fixed #888, #895, #896, #885, #922 Implemented #897 (text output right-to-left).
Fixed issue #888 . Removed ambiguous statements concerning PyMuPDF’s license, which is now clearly stated to be GNU AGPL V3.
Fixed issue #895 .
Fixed issue #896 . Since v1.17.6 PyMuPDF suppresses the font subset tags and only reports the base fontname in text extraction outputs “dict” / “json” / “rawdict” / “rawjson”. Now a new global parameter can request the old behaviour, Tools.set_subset_fontnames() .
Fixed issue #885 . Pixmap creation now also works with filenames given as pathlib.Paths .
Changed Document.subset_fonts() : Text is not rewritten any more and should therefore retain all its original properties – like being hidden or being controlled by Optional Content mechanisms.
Changed TextWriter output to also accept text in right to left mode (Arabian, Hebrew): TextWriter.fill_textbox() , TextWriter.append() . These methods now accept a new boolean parameter right_to_left , which is False by default. Implements #897 .
Changed TextWriter.fill_textbox() to return all lines of text, that did not fit in the given rectangle. Also changed the default of the warn parameter to no longer print a warning message in overflow situations.
Added a utility function recover_quad() , which computes the quadrilateral of a span. This function can be used for correctly marking text extracted with the “dict” or “rawdict” options of Page.get_text() .
This is a bug fix version only. We are publishing early because of the potentially widely used functions.
This is a bug fix version only. We are publishing early because of the potentially widely used functions.
Fixed issue #881 . Fixed a memory leak in Page.insert_image() when inserting images from files or memory.
Fixed issue #878 . pathlib.Path objects should now correctly handle file path hierarchies.
Implemented enhancement requests:
Fixes:
Implemented enhancement requests:
#855, which allows font subsetting using package fontTools
#870, which allows convert_to_pdf method also for PDF documents.
#843, Document.tobytes() (formerly Document.write()) now also support linearized output. Plus several extensions / improvements around supporting Python fileobjects.
Added new methods to quickly determine whether a PDF has annotations or links.
Extended the Document.scrub() method with a new parameter, which allows to also remove page thumbnails.
Added methods to directly inquire and set values in PDF objects - without the need to manipulating PDF object sources in an unwieldy way - see methods Document.xref_set_key() / Document.xref_get_key().
Continued the process of changing the naming convention for class methods and attributes to "snake_case". As announced before, this is a tedious, error-prone process, and requires special care to maintain a high backlevel support for existing scripts.
In future versions - probably synchronously to MuPDF v1.19.0 - we will remove definitions of old names, but a method for re-activating old aliases will remain available.
Added an experimental Document.subset_fonts() which reduces the size of eligible fonts based on their use by text in the PDF. Implements #855 .
Implemented request #870 : Document.convert_to_pdf() now also supports PDF documents.
Renamed Document.write to Document.tobytes() for greater clarity. But the deprecated name remains available for some time.
Implemented request #843 : Document.tobytes() now supports linearized PDF output. Document.save() now also supports writing to Python file objects . In addition, the open function now also supports Python file objects.
Fixed issue #844 .
Fixed issue #838 .
Fixed issue #823 . More logic for better support of OCRed text output (Tesseract, ABBYY).
Fixed issue #818 .
Fixed issue #814 .
Added Document.get_page_labels() which returns a list of page label definitions of a PDF.
Added Document.has_annots() and Document.has_links() to check whether these object types are present anywhere in a PDF.
Added expert low-level functions to simplify inquiry and modification of PDF object sources: Document.xref_get_keys() lists the keys of object xref , Document.xref_get_key() returns type and content of a key, and Document.xref_set_key() modifies the key’s value.
Added parameter thumbnails to Document.scrub() to also allow removing page thumbnail images.
Improved documentation for how to add valid text marker annotations for non-horizontal text.
We continued the process of renaming methods and properties from “mixedCase” to “snake_case” . Documentation usually mentions the new names only, but old, deprecated names remain available for some time.
The recent introduction of "Discussions" by Github has been very motivating for our users. Based on their feedback, several enhancement have been impl
The recent introduction of "Discussions" by Github has been very motivating for our users. Based on their feedback, several enhancement have been implemented. Here is a selection:
There also is a number of fixes - please consult the documentation.
Fixed issue #812 .
Fixed issue #793 . Invalid document metadata previously prevented opening some documents at all. This error has been removed.
Fixed issue #792 . Text search and text extraction will make no rectangle containment checks at all if the default clip=None is used.
Fixed issue #785 .
Fixed issue #780 . Corrected a parameter check error.
Fixed issue #779 . Fixed typo
Added an option to set the desired line height for text boxes. Implements #804 .
Changed text position retrieval to better cope with Tesseract’s glyphless font. Implements #803 .
Added an option to choose the prefix of new annotations, fields and links for providing unique annotation ids. Implements request #807 .
Added getting and setting color and text properties for Table of Contents items for PDFs. Implements #779 .
Added PDF page label handling: Page.get_label() returns the page label, Document.get_page_numbers() return all page numbers having a specified label, and Document.set_page_labels() adds or updates a PDF’s page label definition.
Note
This version introduces Python type hinting . The goal is to provide each parameter and the return value of all functions and methods with type information. This still is work in progress although the majority of functions has already been handled.
Font metrics handling has been improved: text box writing now observes the relevant font properties when determining line heights. In this course a ne
Font metrics handling has been improved: text box writing now observes the relevant font properties when determining line heights. In this course a new option has been introduced, which allows getting text bboxes (glyphs, spans, text search quads, etc.) that more exactly wrap the text only - as opposed to always returning line height bboxes.
Fixes:
Apart from several fixes, this version also focusses on several minor, but important feature improvements. Among the latter is a more precise computation of proper line heights and insertion points for writing / inserting text. As opposed to using font-agnostic constants, these values are now taken from the font’s properties.
Also note that this is the first version which does no longer provide pregenerated wheels for Python versions older than 3.6. PIP also discontinues support for these by end of this year 2020.
Fixed issue #771 . By using “small glyph heights” option, the full page text can be extracted.
Fixed issue #768 .
Fixed issue #750 .
Fixed issue #739 . The “dict”, “rawdict” and corresponding JSON output variants now have two new span keys: "ascender" and "descender" . These floats represent special font properties which can be used to compute bboxes of spans or characters of exactly fontsize height (as opposed to the default line height). An example algorithm is shown in section “Span Dictionary” here . Also improved the detection and correction of ill-specified ascender / descender values encountered in some fonts.
Added a new, experimental Tools.set_small_glyph_heights() – also in response to issue #739 . This method sets or unsets a global parameter to always compute bboxes with fontsize height . If “on”, text searching and all text extractions will returned rectangles, bboxes and quads with a smaller height.
Fixed issue #728 .
Changed fill color logic of ‘Polyline’ annotations: this parameter now only pertains to line end symbols – the annotation itself can no longer have a fill color. Also addresses issue #727 .
Changed Page.getImageBbox() to also compute the bbox if the image is contained in an XObject.
Changed Shape.insertTextbox() , resp. Page.insertTextbox() , resp. TextWriter.fillTextbox() to respect font’s properties “ascender” / “descender” when computing line height and insertion point. This should no longer lead to line overlaps for multi-line output. These methods used to ignore font specifics and used constant values instead.
Improved PDF Optional Content support
Improved PDF Optional Content support
Started overhaul of method and attribute naming
Introduced support of Popup annotations
Implemented the following fixes:
This version adds several features to support PDF Optional Content. Among other things, this includes OCMDs (Optional Content Membership Dictionaries) with the full scope of “visibility expressions” (PDF key /VE ), text insertions (including the TextWriter class) and drawings.
Fixed issue #727 . Freetext annotations now support an uncolored rectangle when fill_color=None .
Fixed issue #726 . UTF-8 encoding errors are now handled for HTML / XML Page.getText() output.
Fixed issue #724 . Empty values are no longer stored in the PDF /Info metadata dictionary.
Added new methods Document.set_oc() and Document.get_oc() to set or get optional content references for existing image and form XObjects. These methods are similar to the same-named methods of Annot .
Added Document.set_ocmd() , Document.get_ocmd() for handling OCMDs.
Added Optional Content support for text insertion and drawing.
Added new method Page.deleteWidget() , which deletes a form field from a page. This is analogous to deleting annotations.
Added support for Popup annotations. This includes defining the Popup rectangle and setting the Popup to open or closed. Methods / attributes Annot.set_popup() , Annot.set_open() , Annot.has_popup , Annot.is_open , Annot.popup_rect , Annot.popup_xref .
Other changes:
The naming of methods and attributes in PyMuPDF is far from being satisfactory: we have CamelCases , mixedCases and lower_case_with_underscores all over the place. With the Annot as the first candidate, we have started an activity to clean this up step by step, converting to lower case with underscores for methods and attributes while keeping UPPERCASE for the constants.
Old names will remain available to prevent code breaks, but they will no longer be mentioned in the documentation.
New methods and attributes of all classes will be named according to the new standard.
As a major new feature, the PDF Optional Content concept is now widely supported.
As a major new feature, the PDF Optional Content concept is now widely supported.
The following fixes have been implemented:
As a major new feature, this version introduces support for PDF’s Optional Content concept.
Fixed issue #714 .
Fixed issue #711 .
Fixed issue #707 : if a PDF user password, but no owner password is supplied nor present, then the user password is also used as the owner password.
Fixed expand and deflate parameters of methods Document.save() and Document.write() . Individual image and font compression should now finally work. Addresses issue #713 .
Added a support of PDF optional content. This includes several new Document methods for inquiring and setting optional content status and adding optional content configurations and groups. In addition, images, form XObjects and annotations now can be bound to optional content specifications. Resolved issue #709 .
and removes the _hit_max_ parameter from text searching. In addition, _hyphenated words_ around line breaks are still found.
This resolves
and removes the hit_max parameter from text searching. In addition, hyphenated words around line breaks are still found.
The use of the clip parameter in text searches and text extractions now only includes characters whose bboxes are fully contained in the clip rctangle.
This version contains some interesting improvements for text searching: any number of search hits is now returned and the hit_max parameter was removed. The new clip parameter in addition allows to restrict the search area. Searching now detects hyphenations at line breaks and accordingly finds hyphenated words.
Fixed issue #575 : if using quads=False in text searching, then overlapping rectangles on the same line are joined. Previously, parts of the search string, which belonged to different “marked content” items, each generated their own rectangle – just as if occurring on separate lines.
Added Document.isRepaired , which is true if the PDF was repaired on open.
Added Document.setXmlMetadata() which either updates or creates PDF XML metadata. Implements issue #691 .
Added Document.getXmlMetadata() returns PDF XML metadata.
Changed creation of PDF documents: they will now always carry a PDF identification ( /ID field) in the document trailer. Implements issue #691 .
Changed Page.searchFor() : a new parameter clip is accepted to restrict the search to this rectangle. Correspondingly, the attribute TextPage.rect is now respected by TextPage.search() .
Changed parameter hit_max in Page.searchFor() and TextPage.search() is now obsolete: methods will return all hits.
Changed character selection criteria in Page.getText() : a character is now considered to be part of a clip if its bbox is fully contained. Before this, a non-empty intersection was sufficient.
Changed Document.scrub() to support a new option redact_images . This addresses issue #697 .
Added transparency options for various methods in classes Shape and Page.
Shape and Page.Fixed issue #692 . PyMuPDF now detects and recovers from more cyclic resource dependencies in PDF pages and for the first time reports them in the MuPDF warnings store.
Fixed issue #686 .
Added opacity options for the Shape class: Stroke and fill colors can now be set to some transparency value. This means that all Page draw methods, methods Page.insertText() , Page.insertTextbox() , Shape.finish() , Shape.insertText() , and Shape.insertTextbox() support two new parameters: stroke_opacity and fill_opacity .
Added new parameter mask to Page.insertImage() for optionally providing an external image mask. Resolves issue #685 .
Added Annot.soundGet() for extracting the sound of an audio annotation.
This version fixes the following issues:
This version fixes the following issues:
Page.cleanContents() should no longer destroy the PDF page's appearance. In earlier versions, this upstream bug occurred in rare cases.Document.insertPDF.The following new features or improvements are included:
Page.getText() now also works for annotations: Annot.getText().Page.getTextbox(rect). This may obsolete extra scripts in many cases.Pixmap.xres and Pixmap.yres values.This is the first PyMuPDF version supporting MuPDF v1.18. The focus here is on extending PyMuPDF’s own functionality – apart from bug fixing. Subsequent PyMuPDF patches may address features new in MuPDF.
Fixed issue #519 . This upstream bug occurred occasionally for some pages only and seems to be fixed now: page layout should no longer be ruined in these cases.
Fixed issue #675 .
Unsuccessful storage allocations should now always lead to exceptions (circumvention of an upstream bug intermittently crashing the interpreter).
Pixmap size is now based on size_t instead of int in C and should be correct even for extremely large pixmaps.
Fixed issue #668 . Specification of dashes for PDF drawing insertion should now correctly reflect the PDF spec.
Fixed issue #669 . A major source of memory leakage in Page.insert_pdf() has been removed.
Added keyword “images” to Page.apply_redactions() for fine-controlling the handling of images.
Added Annot.getText() and Annot.getTextbox() , which offer the same functionality as the Page versions.
Added key “number” to the block dictionaries of Page.getText() / Annot.getText() for options “dict” and “rawdict”.
Added glyph_name_to_unicode() and unicode_to_glyph_name() . Both functions do not really connect to a specific font and are now independently available, too. The data are now based on the Adobe Glyph List .
Added convenience functions adobe_glyph_names() and adobe_glyph_unicodes() which return the respective available data.
Added Page.getDrawings() which returns details of drawing operations on a document page. Works for all document types.
Improved performance of Document.insert_pdf() . Multiple object copies are now also suppressed across multiple separate insertions from the same source. This saves time, memory and target file size. Previously this mechanism was only active within each single method execution. The feature can also be suppressed with the new method bool parameter final=1 , which is the default.
For PNG images created from pixmaps, the resolution (dpi) is now automatically set from the respective Pixmap.xres and Pixmap.yres values.
Fixed #651 Fixed #645 Fixed #622 Fixed #653 Fixed #640 Added methods and atrributes to speed up TOC maintenance. Added new page method to extract text
Fixed #651
Fixed #645
Fixed #622
Fixed #653
Fixed #640
Added methods and atrributes to speed up TOC maintenance.
Added new page method to extract text from inside a rectangle.
All getText() methods (except (X)HTML and XML) now support a clip parameter.
Fixed issue #651 . An upstream bug causing interpreter crashes in corner case redaction processings was fixed by backporting MuPDF changes from their development repo.
Fixed issue #645 . Pixmap top-left coordinates can be set (again) by their own method, Pixmap.set_origin() .
Fixed issue #622 . Page.insertImage() again accepts a rect_like parameter.
Added several new methods to improve and speed-up table of contents (TOC) handling. Among other things, TOC items can now changed or deleted individually – without always replacing the complete TOC. Furthermore, access to some PDF page attributes is now possible without first loading the page. This has a very significant impact on the performance of TOC manipulation.
Added an option to Document.insert_pdf() which allows displaying progress messages. Addresses #640 .
Added Page.getTextbox() which extracts text contained in a rectangle. In many cases, this should obsolete writing your own script for this type of thing.
Added new clip parameter to Page.getText() to simplify and speed up text extraction of page sub areas.
Added TextWriter.appendv() to add text in vertical write mode . Addresses issue #653
Added origin key to text span dictionary of Page.getText("dict").
origin key to text span dictionary of Page.getText("dict").buffer to fitz.Font.sanitize to Page.cleanContents().Fixed issue #605
Fixed issue #600 – text should now be correctly positioned also for pages with a CropBox smaller than MediaBox.
Added text span dictionary key origin which contains the lower left coordinate of the first character in that span.
Added attribute Font.buffer , a bytes copy of the font file.
Added parameter sanitize to Page.cleanContents() . Allows switching of sanitization, so only syntax cleaning will be done.
Correct use of opacity default in TextWriter.writeText().
opacity default in TextWriter.writeText().Fixed issue #561 – second go: certain TextWriter usages with many alternating fonts did not work correctly.
Fixed issue #566 .
Fixed issue #568 .
Fixed – opacity is now correctly taken from the TextWriter object, if not given in TextWriter.writeText() .
Added a new global attribute fitz_fontdescriptors . Contains information about usable fonts from repository pymupdf-fonts .
Added Font.valid_codepoints() which returns an array of unicode codepoints for which the font has a glyph.
Added option text_as_path to Page.getSVGimage() . this implements #580 . Generates much smaller SVG files with parseable text if set to False .
Fixes #561 - more than 10 TextWriter fonts per page Fixes #562 - annotation pixmaps no longer derived from tha page Fixes wrong appearance of mono-spa
Fixes #561 - more than 10 TextWriter fonts per page
Fixes #562 - annotation pixmaps no longer derived from tha page
Fixes wrong appearance of mono-spaced fonts
Resolves #563 - allow manipulation of PDF property NeedAppearances
Support optional fonts provided via repository pymupdf-fonts
Fixed issue #561 . Handling of more than 10 Font objects on one page should now work correctly.
Fixed issue #562 . Annotation pixmaps are no longer derived from the page pixmap, thus avoiding unintended inclusion of page content.
Fixed issue #559 . This MuPDF bug is being temporarily fixed with a pre-version of MuPDF’s next release.
Added utility function repair_mono_font() for correcting displayed character spacing for some mono-spaced fonts.
Added utility method Document.need_appearances() for fine-controlling Form PDF behavior. Addresses issue #563 .
Added utility function sRGB_to_pdf() to recover the PDF color triple for a given color integer in sRGB format.
Added utility function sRGB_to_rgb() to recover the (R, G, B) color triple for a given color integer in sRGB format.
Added utility function make_table() which delivers table cells for a given rectangle and desired numbers of columns and rows.
Added support for optional fonts in repository pymupdf-fonts .
Fixes #540 Fixes #548 TextWriter.fillTextbox now supports indenting start of text.
Fixes #540
Fixes #548
TextWriter.fillTextbox now supports indenting start of text.
Fixed an undocumented issue, which prevented fully cleaning a PDF page when using Page.cleanContents() .
Fixed issue #540 . Text extraction for EPUB should again work correctly.
Fixed issue #548 . Documentation now includes LINK_NAMED .
Added new parameter to control start of text in TextWriter.fillTextbox() . Implements #549 .
Changed documentation of Page.add_redact_annot() to explain the usage of non-builtin fonts.
Fixes #533 Adds appearance flexibility to redactions annotations, #535.
Fixes #533 Adds appearance flexibility to redactions annotations, #535.
Fixed issue #533 .
Added options to modify ‘Redact’ annotation appearance. Implements #535 .
Fixed: #525, #520 Implemented: #524
Fixed: #525, #520 Implemented: #524
Improved interactive help a lot.
Fixed issue #520 .
Fixed issue #525 . Vertices for ‘Ink’ annots should now be correct.
Fixed issue #524 . It is now possible to query and set rotation for applicable annotation types.
Also significantly improved inline documentation for better support of interactive help.
Improved Redaction annotation support.
This version is based on MuPDF v1.17. Following are highlights of new and changed features:
Added extended language support for annotations and widgets: a mixture of Latin, Greece, Russian, Chinese, Japanese and Korean characters can now be used in ‘FreeText’ annotations and text widgets. No special arrangement is required to use it.
Faster page access is implemented for documents supporting a “chapter” structure. This applies to EPUB documents currently. This comes with several new Document methods and changes for Document.loadPage() and the “indexed” page access doc[n] : In addition to specifying a page number as before, a tuple (chapter, pno) can be specified to identify the desired page.
Changed: Improved support of redaction annotations: images overlapped by redactions are permanently modified by erasing the overlap areas. Also links are removed if overlapped by redactions. This is now fully in sync with PDF specifications.
Other changes:
Changed TextWriter.writeText() to support the “morph” parameter.
Added methods Rect.morph() , IRect.morph() , and Quad.morph() , which return a new Quad .
Changed Page.add_freetext_annot() to support text alignment via a new “align” parameter.
Fixed issue #508 . Improved image rectangle calculation to hopefully deliver correct values in most if not all cases.
Fixed issue #502 .
Fixed issue #500 . Document.convertToPDF() should no longer cause memory leaks.
Fixed issue #496 . Annotations and widgets / fields are now added or modified using the coordinates of the unrotated page . This behavior is now in sync with other methods modifying PDF pages.
Added Page.rotationMatrix and Page.derotationMatrix to support coordinate transformations between the rotated and the original versions of a PDF page.
Potential code breaking changes:
The private method Page._getTransformation() has been removed. Use the public Page.transformationMattrix instead.
This introduces two new classes (Font and TextWriter) for better support of MuPDF's text writing features for PDF files. In addition, issues #488 and
This introduces two new classes (Font and TextWriter) for better support of MuPDF's text writing features for PDF files. In addition, issues #488 and #493 are fixed.
Note: This should be the last version based on MuPDF v1.16. Artifex has just published (April 21) a release candidate for a new version 1.7.
This version introduces several new features around PDF text output. The motivation is to simplify this task, while at the same time offering extending features.
One major achievement is using MuPDF’s capabilities to dynamically choosing fallback fonts whenever a character cannot be found in the current one. This seamlessly works for Base-14 fonts in combination with CJK fonts (China, Japan, Korea). So a text may contain any combination of characters from the Latin, Greek, Russian, Chinese, Japanese and Korean languages.
Fixed issue #493 . Pixmap(doc, xref) should now again correctly resemble the loaded image object.
Fixed issue #488 . Widget names are now modifiable.
Added new class Font which represents a font.
Added new class TextWriter which serves as a container for text to be written on a page.
Added Page.writeText() to write one or more TextWriter objects to the page.
- Fixed issue #479 . PyMuPDF should now more correctly report image resolutions. This applies to both, images (either from images files or extracted f
Fixed issue #479 . PyMuPDF should now more correctly report image resolutions. This applies to both, images (either from images files or extracted from PDF documents) and pixmaps created from images.
Added Pixmap.set_dpi() which sets the image resolution in x and y directions.
Fixes #477 and #476. Also fixes an error which prevented an empty interior for 'Polyline' and 'Polygon' annotations, when the stroke color was set. In
Fixes #477 and #476.
Also fixes an error which prevented an empty interior for 'Polyline' and 'Polygon' annotations, when the stroke color was set.
In addition, the interior of line end symbols (where applicable) is now either white, or the interior of the annotation, or an independent fill color set by Annot.update parameter fill_color. Line ends without an interior area remain unchanged (e.g. open arrows, butt, etc.).
Fixed issue #477 .
Fixed issue #476 .
Changed annotation line end symbol coloring and fixed an error coloring the interior of ‘Polyline’ /’Polygon’ annotations.
Nothing published for this version
Fixes for issues: #474 (endcoding errors) #473 clarify documentation of name changes #467 introduce manipulation of anti-aliasing levels #466 ensure c
Fixes for issues:
#474 (endcoding errors)
#473 clarify documentation of name changes
#467 introduce manipulation of anti-aliasing levels
#466 ensure correct of use of unicode type
#453 introduce new Document method for PDF scrubbing (sensitive information removal)
#416 allow text line highlighting between two given points and introduce PDF 'blend mode' for annotations
Changed text marker annotations to accept parameters beyond just quadrilaterals such that now text lines between two given points can be marked .
Added Document.scrub() which removes potentially sensitive data from a PDF. Implements #453 .
Added Annot.blendMode() which returns the blend mode of annotations.
Added Annot.setBlendMode() to set the annotation’s blend mode. This resolves issue #416 .
Changed Annot.update() to accept additional parameters for setting blend mode and opacity.
Added advanced graphics features to control the anti-aliasing values , Tools.set_aa_level() . Resolves #467
Fixed issue #474 .
Fixed issue #466 .
Fixes #465 and implements #464.
Fixes #465 and implements #464.
Added Document.getPageXObjectList() which returns a list of Form XObjects of the page.
Added Page.setMediaBox() for changing the physical PDF page size.
Added Page methods which have been internal before: Page.cleanContents() (= Page._cleanContents() ), Page.getContents() (= Page._getContents() ), Page.getTransformation() (= Page._getTransformation() ).
A number of issues are also being fixed with this release:
A number of issues are also being fixed with this release: #447, #461, #397, #463, #454, #434
Fixed issue #447
Fixed issue #461 .
Fixed issue #397 .
Fixed issue #463 .
Added JavaScript support to PDF form fields, thereby fixing #454 .
Added a new annotation method Annot.delete_responses() , which removes ‘Popup’ and response annotations referring to the current one. Mainly serves data protection purposes.
Added a new form field method Widget.reset() , which resets the field value to its default.
Changed and extended handling of redactions: images and XObjects are removed if contained in a redaction rectangle. Any partial only overlaps will just be covered by the redaction background color. Now an overlay text can be specified to be inserted in the rectangle area to take the place the deleted original text. This resolves #434 .
This release fixes issues #344, #426, #434, #443, #444.
This release fixes issues #344, #426, #434, #443, #444.
Added Support for redaction annotations via method Page.add_redact_annot() and Page.apply_redactions() .
Fixed issue #426 (“PolygonAnnotation in 1.16.10 version”).
Fixed documentation only issues #443 and #444 .
Fixing issues #415, #417 and #421.
Fixing issues #415, #417 and #421.
Fixed issue #421 (“annot.set_rect(rect) has no effect on text Annotation”)
Fixed issue #417 (“Strange behavior for page.deleteAnnot on 1.16.9 compare to 1.13.20”)
Fixed issue #415 (“Annot.setOpacity throws mupdf warnings”)
Changed all “add annotation / widget” methods to store a unique name in the /NM PDF key.
Changed Annot.setInfo() to also accept direct parameters in addition to a dictionary.
Changed Annot.info to now also show the annotation’s unique id ( /NM PDF key) if present.
Added Page.annot_names() which returns a list of all annotation names ( /NM keys).
Added Page.load_annot() which loads an annotation given its unique id ( /NM key).
Added Document.reload_page() which provides a new copy of a page after finishing any pending updates to it.
Your coding agent can read these notes before it upgrades. Set up the MCP server →