NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1575 most downloaded on PyPI
SDK and CLI for parsing PDF, DOCX, HTML, and more, to a unified document representation for powering downstream workflows such as gen AI applications.
Last release 3 days ago
18 Sep 2026
Ships on a steady schedule
a new release about every 9 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
217 releases · first in 2024
One column per month.
Add tutorial using Milvus and Docling for RAG pipeline (#1449) (`a2fbbba`)
ed20124)8012a3e)fa7fc9e)c2470ed)64918a8)995b3b0)88948b0)a7dd59c)293c28c)01fbfd5)a026b4e)cli: Add option for html with split-page mode (#1355) (`c0ba88e`)
c0ba88e)eef2bde)c605edd)7e40ad3)0de70e7)415b877)2503999)6b696b5)Handle tags as code blocks (#1320) (`0499cd1`)
<code> tags as code blocks (#1320) (0499cd1)bfcab3d)14e9c0c)dc3bf9c)d2d6874)b3d111a)Fixes tables when using OCR (#1261) (`7afad7e`)
Word-level pdf cells for tables (#1238) (`8bd71e8`)
Improve HTML layer detection, various MD fixes (#1241) (`9210812`)
converter: Cache same pipeline class with different options (#1152) (`825b226`)
SmolDocling: Support MLX acceleration in VLM pipeline (#1199) (`1c26769`)
1c26769)b454aa1)2f72167)f5adfb9)0b707d0)Add factory for ocr engines via plugins (#1010) (`6eaae3c`)
6eaae3c)3960b19)772487f)6eb718f)f94da44)0945973)aa92a57)Use new TableFormer model weights and default to accurate model version (#1100) (`eb97357`)
Proper handling of orphan IDs in layout postprocessing (#1118) (`c56ab3a`)
Enable locks for threadsafe pdfium (#1052) (`8dc0562`)
[Experimental] Introduce VLM pipeline using HF AutoModelForVision2Seq, featuring SmolDocling model (#1054) (`3c9fe76`)
3c9fe76)ab683e4)e197225)1b0ead6)Implement new reading-order model (#916) (`c93e369`)
Runtime error when Pandas Series is not always of string type (#1024) (`6796f0a`)
Support cuda:n GPU device allocation (#694) (`77eb77b`)
Add support for CSV input with new backend to transform CSV files to DoclingDocument (#945) (`00d9405`)
00d9405)2716c7d)5101e25)af19c03)c47ae70)Add content_layer property to items to address body, furniture and other roles (#735) (`cf78d5b`)
Describe pictures using vision models (#259) (`4cc6e3e`)
New artifacts path and CLI utility (#876) (`ed74fe2`)
90b766e)9114ada)722a6eb)5ad6de0)Expose equation exports (#869) (`6a76b49`)
6a76b49)70d68b6)d727b04)4df085a)5ac2887)94751a7)0cd81a8)eff16b6)b1cf796)2c037ae)2a1f8af)bccb022)fea0a99)CLI: Expose code and formula models in the CLI (#820) (`6882e6c`)
6882e6c)95b293a)rec_keys_path in RapidOcrOptions to support custom dictionaries (#786) (5332755)3be2fb5)5aed9f8)adf6353)a112d7a)6875913)5139b48)4d41db3)8a4ec77)b885b2f)c2ae1cc)New document picture classifier (#805) (`16a218d`)
16a218d)88a0e66)3213b24)8543c22)a458e29)670a08b)Improve OCR results, stricten criteria before dropping bitmap areas (#719) (`5a060f2`)
Added http header support for document converter and cli (#642) (`0ee849e`)
Create a backend to transform PubMed XML files to DoclingDocument (#557) (`fd03480`)
Updated Layout processing with forms and key-value areas (#530) (`60dc852`)
Introduce support for GPU Accelerators (#593) (`19fad92`)
Add timeout limit to document parsing job. DS4SD#270 (#552) (`3da166e`)
Docling-parse v2 as default PDF backend (#549) (`aca57f0`)
Expose new hybrid chunker, update docs (#384) (`c8ecdd9`)
c8ecdd9)3e073df)eb7ffcd)py.typed marker file (#531) (9102fe1)0d11e30)b730b2d)c830b92)8ada0bc)Improve handling of disallowed formats (#429) (`34c7c79`)
ParserError EOF inside string (#470) (#472) (`c90c41c`)
c90c41c)d3f84b2)5ba3807)33cff98)d487210)8ccb3c6)Added support for exporting DocItem to an image when page image is available (#379) (`3f91e7d`)
3f91e7d)ed785ea)926dfd2)7a97d71)8533039)8b437ad)Skip glm model downloads (#322) (`c9341bf`)
OCR: Introduce the OcrOptions.force_full_page_ocr parameter that forces a full page OCR scanning (#290) (`c6b3763`)
c6b3763)5d4a10b)81c8243)97f214e)EasyOcrModel: Support the use_gpu pipeline parameter in EasyOcrModel. Initialize easyocr (#282) (`0eb065e`)
tesserocr: Raise Exception if tesserocr has not loaded any languages (#279) (`704d792`)
Pdf backend, table mode as options and artifacts path (#203) (`40ad987`)
Simplify torch dependencies and update pinned docling deps (#190) (`eb679cc`)
Add pipeline timings and toggle visualization, establish debug settings (#183) (`2a2c65b`)
Fix header levels for DOCX & HTML (#184) (`b9f5c74`)
b9f5c74)94d0729)7d19418)88c1673)Add coverage_threshold to skip OCR for small images (#161) (`b346faf`)
New experimental docling-parse v2 backend (#131) (`5e4944f`)
Remove stderr from tesseract cli and introduce fuzziness in the text validation of OCR tests (#138) (`dae2a3b`)
Add options for choosing OCR engines (#118) (`f96ea86`)
New torch-based docling models (#120) (`2422f70`)
Windows support (#122) (`d44c62d`)
Allow usage of opencv 4.6.x (#110) (`34bd887`)
Support tableformer model choice (#90) (`d6df76f`)
Add figure in markdown (#98) (`6a03c20`)
Updated the render_as_doctags with the new arguments from docling-core (#93) (`4794ce4`)
Your coding agent can read these notes before it upgrades. Set up the MCP server →