NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2121 most downloaded on PyPI
Simple package to extract text with coordinates from programmatic PDFs
Last release 6 days ago
28 Sep 2026
Ships fairly regularly
a new release about every 9 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
2 years old
98 releases · first in 2024
One column per month.
Updated the page_item_sanitators/cells.h and its associated ground-truth
Bbox of brackets, summations, missing font-mappings etc
Answer shape geometry queries in the frame of the text cells
Stray PostScript resource directives no longer poison the next colour operator
Adding static page-count methods
Resolve-bookmarks-using-page-aware-PDF-outlines
Optimization of the parse/render with up to 3.84× speedup
Regression-test overhaul, and the parser/renderer fixes it uncovered
render: Clip masks, shading patterns and Coons meshes; CCITT, CMYK-JPEG, tiling-pattern and CJK text fixes
render: Clip paths, masks, CMYK/Indexed images, font fallback, Symbol brackets
render: Add tiling patterns, clipping, Type3 glyphs and CID text
Make Blend2D font fallback portable across distros
Improve PDF rendering fidelity across color, image, transparency, and writing-mode cases
Added the rendering-regression with pypdfium
Honor /ActualText replacement text of marked-content spans
Report document load failures instead of deferring them as -1 page count
Adding new top-level methods for page content geometry queries
Fixing the word extraction with spaces as blockers (#296) (`66e4c33`)
Build wheel for windows arm64 (#288) (`9acdc48`)
Working on refactoring the parse for shapes (#293) (`cbe73a8`)
Attack embedded fonts (#291) (`f41aec4`)
Reduce total number of results cached (#287) (`55d8476`)
Caching objects (#285) (`2bc547e`)
Upgrade bitmap rendering with clipping (#284) (`b55da56`)
Improve pages/sec and scalability, new decode configuration shape (#278) (`1d8f70c`)
Add materialize_bitmap_bytes flag to skip bitmap byte extraction (#277) (`918d7b9`)
Release unused python memory (#274) (`537461a`)
Upgrade packages for vulnerabilities (#270) (`1515795`)
Rendering of math and latex symbols (#264) (`ac0a361`)
Memory management for docling upstream (#263) (`db84017`)
Adding the cpp analysis script and enhancing the extraction of bitmap types (fix for rotated images). (#250) (`70fa300`)
Improve extraction from fillable fields (#247) (`c3c1e85`)
Extend the renderer (#245) (`e7ef57f`)
Prevent infinite loop in TOC extraction with circular PDF refererences (#246) (`092d1b8`)
Bo10k document failures (#244) (`1f650dd`)
Add parallelization for parsing (#216) (`ae66f6d`)
Ligatures and unicode chars in Differences (#234) (`856c0fe`)
Map characters into the proper chars (#233) (`0316060`)
Add config option to remove glyph output (#231) (`9657023`)
Robustify parse of broken pdfs (#228) (`e0264dd`)
Replace fixed-size utf8::append buffers with std::back_inserter to prevent segfaults (#224) (`237cef6`)
Rotated pages (missing commits) (#219) (`6d98479`)
Deal with image containing rotated pages (#217) (`0b592f6`)
Refactor pdf resources to pdf page item (#215) (`e7812a1`)
e7812a1)67d2922)3272dd8)ea5f1d8)f01ce84)25672da)Add typed serialization (#201) (`23c7fb8`)
Remove python3.9 (#200) (`d162c32`)
Remove deprecated v1 api (#189) (`adcb9b0`)
"could not find the page-dimensions" error solved restoring the parent mediabox (#181) (`1d3f78e`)
360 rotated pages (#177) (`327dc4b`)
Support reading password protected PDF (#169) (`0c64402`)
Support for py3.14 (#174) (`5caf1ff`)
Support pdf with only trim-bbox (#173) (`76ab6b5`)
Add perf tools (#165) (`f8d53ee`)
Reset to the old parameters in sanitation (#163) (`0402b3f`)
Accelerate docling parse (#161) (`1466548`)
Your coding agent can read these notes before it upgrades. Set up the MCP server →