NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1224 most downloaded on PyPI
PyMuPDF Utilities for LLM/RAG
Last release 1 months ago
06 Aug 2026
Release timing varies
gaps range from 8 days to 4 months
Most releases are documented
notes for 34 of 47 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
47 releases · first in 2024
pdf4llm/setup.py setup.py: increment version to 1.28.2.
pdf4llm/setup.py setup.py: increment version to 1.28.2.
Fixed issues:
Fixed 4670 : scrub fails to remove hidden text after clean_contents stopped including line breaks (u2265 1.24.0)
Fixed 4943 : enh: applying redactions with image cropping for currently unsupported colorspaces
Fixed 5030 : find_tables() with layout enabled can return a zero-cell Table, and Table.bbox then raises “ValueError: min() iterable argument is empty”
Fixed 5042 : Page.get_texttrace() leaks None references u2014 Fatal Python error (none_dealloc) in long-running processes (1.27.2.3 & 1.28.0)
Fixed 5044 : Outline (TOC) parsing bug
Fixed 5049 : font subsetting segfaults in 1.28.0, regression from 1.27.2.2 #5049
Fixed 5056 : Option for reproducible/deterministic save output (omit MuPDF version banner + Producer)
Other:
Use MuPDF-1.28.2.
Output warning when legacy fitz module is imported.
Cope better with markdown containing illegal utf8 sequences.
Fixed building with PYMUPDF_SETUP_MUPDF_VS_UPGRADE.
pymupdf.Page.find_tables():
Retrospectively added fix for #4936 in release 1.28.0 below.
Added windows-arm64 wheel.
One column per month.
setup.py: add _build.py file to wheel/install containing git info.
setup.py: add _build.py file to wheel/install containing git info.
Fixed issues:
Fixed 4114 : ComboBox choice_values full of empty strings despite PDF having valid choices.
Fixed 4950 : remove_rotation() raises ValueError on widgets with empty/infinite rects
Fixed 5001 : Formulae incorrectly rendered as black boxes
Fixed 5033 : Annot.set_rotation(0) followed by Annot.update() throws AttributeError
Fixed 4936 : bug: incomplete redaction of vector graphics (line art)
Other:
Use MuPDF-1.28.0.
pymupdf.Document.init(): new arg Archive to support documents with archives.
pymupdf.Document.convert_to_pdf(): also generate links.
Fixed MacOS x64 platform tag - used to be macosx_10_9_x86_64, but we actually require MacOS 10.15, so platform tag is now macosx_10_15_x86_64.
Support Windows builds with free thread python.
pymupdf.Document.save() now saves non-PDF documents in PDF format.
New method pymupdf.Document.apply_css().
Pyodide wheel is available on pypi.org.
pdf4llm/setup.py: increment version to 1.27.2.3.
pdf4llm/setup.py: increment version to 1.27.2.3.
Fixed issues:
Fixed 4928 : pymupdf.Document.scrub raises AttributeError for a document with annotations
Fixed 4942 : bug: IndexError for Page.get_links after Page.clip_to_rect
Fixed 4954 : get_drawings() returns incorrect lineJoin and width
Fixed 4958 : bug: inserting rotated pages to another document messes up link coordinates
Other:
Fixed incorrect generation of lineJoin j in PDF content, introduced in 1.27.2.2.
Allow build to (incorrectly) claim to be thread-safe, for #4760. See setup.py for details.
Use pypi.org’s pipcl package instead of our own pipcl.py file.
CHANGES.md: add version number.
CHANGES.md: add version number.
ocr installation folder.force_ocr=True does no longer require to specify ocr_function. If no OCR function is given, the best available plugin is chosen. An exception is raised only if none of the plugins is usable.Fixed issues:
Fixed 4902 : Incorrect linewidth in elements returned by Page.get_texttrace()
Fixed 4932 : “Page” has no attribute “find_tables” in PyMuPDF 1.27
Other:
Added Annot.bool() .
pymupdf4llm/: fix naming bug and update version to 1.27.2.1.
pymupdf4llm/: fix naming bug and update version to 1.27.2.1.
Pymupdf4llm now automatically installs and uses pymupdf_layout.
Installing the pympdf4llm package automatically installs pymupdf_layout.
import pymupdf4llm will automatically initialise layout.
Layout can be disabled by calling pymupdf4llm.use_layout(False).
Our release numbering scheme has been changed to comply with the other packages in the PyMuPDF family.
For change details please consult this file.
For change details please consult this file.
ocr_function=None. If not None it must be a callable, which is expected to OCR the page giving it a text layer.force_ocr=False to all extraction functions. Requires to also provide a callable via ocr_function. If True, ocr_function is called for every page, thus skipping the otherwise executed "OCR worthiness" check.For change details please consult this file.
For change details please consult this file.
For a description of changes see this file.
For a description of changes see this file.
get_key_values() to extract the field name and their values if the document is a "Form PDF". This is always available, whether or not PyMuPDF-Layout is avtive.For a summary of changes see file https://github.com/pymupdf/pymupdf4llm/blob/main/CHANGES.md
For a summary of changes see file https://github.com/pymupdf/pymupdf4llm/blob/main/CHANGES.md
Support new parameter ocr_language. This is a string which is passed through to Tesseract-OCR, so the user is responsible for its format.
Changed the format of the page chunk dictionary: the new dictionary key "page_boxes" in layout mode is now a list of dictionaries (was a list of lists). The dictionaries have the following keys / values:
"index": 0-based integer enumerating the layout boxes in reading order
"class": a string denoting the bbox class ("table", "list-item", "section-header", etc.)
"bbox": pymupdf.IRect of the layout boundary box
"pos": tuple (start, stop) of integers denoting the text substring of the bboxes text in this chunk's text ("text" key of the chunk). The values are in slice format and can be used to extract the bbox text like this bbox_text = chunk["text"][start : stop].
Implemented multiple performance improvements, primarily around rectangle containment checks.
323 - page_chunks=True parameter was ignored in PyMuPDF-Layout mode
page_chunks=True parameter was ignored in PyMuPDF-Layout modeto_markdown() / to_text() now both support Page chunk output via parameter page_chunks=True.Forum - List index out of range ...
341 - Broken markdown parsing for new line directly followed by 'o'...
table_format in method to_text() (PyMuPDF-Layout only). This allows selecting the appearance of tables in plain text outputs. The possible values are defined in the list tabulate.tabulate_formats. Default is "grid".pip command: pip install --upgrade pymupdf4llm[ocr,layout]. This will install pymupdf4llm, pymupdf, and pymupdf-layout. The "ocr" parameter - when needed - installs opencv-python for automatic OCR support in PyMuPDF-Layout mode. Combine this with parameters --upgrade, --force-reinstall or --no-cache-dir as necessary.### Fixes: * 335 - KeyError "has_ocr_text" ### Other Changes: ------
332 - TypeError("to_markdown() got an unexpected keyword argument 'header'")
ocr_dpi=400 which sets the OCR resolution for full-page OCR.StructTreeRoot objects in PDF.NotImplementedError exceptions when layout-only features are used.page_separators as in the legacy mode.Nothing published for this version
320 - [Bug] ValueError: min() iterable argument is empty ...
This version introduces full support of the PyMuPDF-Layout package. This entails a radically new approach for detecting the layout of document pages u
This version introduces full support of the PyMuPDF-Layout package. This entails a radically new approach for detecting the layout of document pages using the AI-based features of the layout package.
Improvements include:
The PyMuPDF-Layout package is not open-source and has its own license, which is different from PyMuPDF4LLM. It also is dependent on a number of other, fairly large packages like onnxruntime, numpy, sympy and OpenCV, which each in turn have their own dependencies.
We therefore keep the use of the layout feature optional. To activate PyMuPDF-Layout support the following import statement must be included before importing PyMuPDF4LLM itself:
import pymupdf.layout
import pymupdf4llm
Thereafter, PyMuPDF's namespace is available. The known method pymupdf4llm.to_markdown() automatically works with AI-based empowerment.
In addition, two new methods become available:
pymupdf4llm.to_text() - which works much like markdown output but produces plain text.pymupdf4llm.to_json() - which outputs the document's metadata and the selected pages in JSON format.show_progress=True, Python package tqdm is automatically used when available to display a progress bar. If tqdm is not installed, our own text-based progress bar is used.Nothing published for this version
Nothing published for this version
Nothing published for this version
296 - [Bug] A specific diagram recognized as significant ...
to_markdown: page_separators=False. If True and page_chunks=False a line like --- end of page=nnn --- is appended to each pages markdown text. The page number is 0-based. Intended for debugging purposes.289 - Content Duplication with the latest version
The table module in package PyMuPDF has been modified: Its method to_markdown() will now output markdown-styled cell text. Previously, table cells were extracted as plain text only.
The class TocHeaders is now a top-level import and can now be directly used.
Method to_markdown has a new parameter detect_bg_color=True (default) which guesses the page's background color. If a background is detected, fill-only vectors having this color are ignored. False will always consider "fill" vectors in vector graphics detection.
Text written with a Type 3 font will now always be considered. Previously, this text was always treated as invisible and was hence suppressed.
The package now contains the license file GNU Affero GPL 3.0 to ease distribution (see LICENSE). It also clarifies that PyMuPDF4LLM is dual licensed under GNU AGPL 3.0 and individual commercial licenses.
There is a new file versions_file.py which contains version information. This is used to ensure the presence of a minimum PyMuPDF version at import time.
282 - Content Duplication with the latest version
The table module in package PyMuDDF has been: Its method to_markdown() will now output markdown-styled cell text. Previously, table cells were extracted as plain text only.
The class TocHeaders is now a top-level import and can now be directly used.
Text written with a Type 3 font will now always be considered. Previously, this text was always treated as invisible and was hence suppressed.
### Fixes: * Fixing "UnboundLocalError" ### Other Changes:
263 - Table Strategy = None raises error
251 - Images a little larger than the page size are being ignored
116 - Handling Graphical Images & Superscripts
171 - Text rects overlap with tables and images that should be excluded.
Added new parameter ignore_images: (bool) optional. True will not consider images in any way. May be useful for pages where a plethora of images prevents meaningful layout analysis. Typical examples are PowerPoint slides and derived / similar pages.
Added new parameter ignore_graphics: (bool), optional. True will not consider graphics except for table detection. May be useful for pages where a plethora of vector graphics prevents meaningful layout analysis. Typical examples are PowerPoint slides and derived / similar pages.
Added new parameter to class IdentifyHeaders: Use max_levels (integer <= 6) to limit the generation of header tag levels. e.g. headers = pymupdf4llm.IdentifyHeaders(doc, max_level=3) ensures that only up to 3 header levels will ever be generated. Any text with a font size less than the value of ### will be body text. In this case, the markdown generation itself would be coded as md = pymupdf4llm.to_markdown(doc, hdr_info=headers, ...).
Changed parameter table_strategy: When specifying None, no effort to detecting tables will be made. This can be useful when tables are of no interest or known to not exist in a given file. This will speed up processing significantly. Be prepared to see more changes and extensions here.
The following list includes fixes made in version 0.0.18 already.
The following list includes fixes made in version 0.0.18 already.
Added new parameter filename: (str), optional. Overwrites or sets the filename for saved images. Useful when the document is opened from memory.
Added new parameter use_glyphs: (bool), optional. Request to use the glyph number (if possible) of a character if the font has no back-translation to the original Unicode value. The default is False which causes � symbols to be rendered in these cases.
Added strike-out support: We now detect and render striked-out text.
Improved background color detection: We have introduced a simple background color detection mechanism: If a page shows an identical color in all four corners, we assume this to be the background color. Text and vector graphics with this color will be ignored as invisible.
Improved invisible text detection: Text with an alpha value of 0 is now ignored.
Improved fake-bold detection: Text mimicking bold appearance is now treated like standard bold text in most cases.
Header handling changes:
Changed handling of parameter graphics_limit: We previously ignored a page completely if the vector graphics count exceeded the limit. We now only ignore vector graphics if their count outside table boundary boxes is too large. This should only suppress vector graphics on the page, while keeping images, text and table content extractable.
Changed the margins default to 0. The previous default (0, 50, 0, 50) ignored 50 points at the top and bottom of pages. This has turned out to cause confusion in too many cases.
Nothing published for this version
147 - Error when page contains nothing but a table.
Nothing published for this version
138 - Table is not extracted and some text order was wrong.
embed_images (bool) embeds images and vector graphics in the markdown text as base64-encoded strings. Ignores write_images and image_path parameters.image_size_limit which is a float between 0 and 1, default is 0.05 (5%). Causes images to be ignored if their width or height values are smaller than the corresponding fraction of the page's width or height.Nothing published for this version
112 - Invalid bandwriter header dimensions/setup.
ignore_code suppresses special formatting of text in mono-spaced fonts.extract_words enforces page_chunks=True and adds a "words" list to each page dictionary.Nothing published for this version
90 - 'Quad' object has no attribute 'tl'.
73 - bug in to_markdown internal function.
to_markdown internal function.image_format.image_path.write_images=True, text on images / graphics can be suppressed by setting force_text=False.71 - Unexpected results in pymupdf4llm but pymupdf works.
65 - Fix typo in pymupdf_rag.py.
pymupdf_rag.py.54 - Mistakes in orchestrating sentences. Additional fix: text extraction no longer uses the TEXT_DEHYPHENATE flag bit.
TEXT_DEHYPHENATE flag bit.55 - Bug in helpers/multi_column.py - IndexError: list index out of range.
dpi to specify the resolution of images.page_width / page_height for easily processing reflowable documents (Text, Office, e-books).graphics_limit to avoid spending runtimes for value-less content.table_strategy to directly control the table detection strategy.Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →