NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3756 most downloaded on PyPI
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Last release 6 days ago
28 Sep 2026
Ships fairly regularly
a new release about every 3 weeks
Some releases are documented
notes for 15 of the last 60 stable releases
3 versions withdrawn
withdrawn after publishing
11 years old
271 releases · first in 2015
PDF/A made without Ghostscript ("speculative" conversion) is now validated with pikepdf's PDF/A support ( pikepdf.pdfa ) instead of veraPDF, and the f
Changes
pikepdf.pdfa) instead of veraPDF, and the-v1/CIDSet PDF/A-1 requires, and removes XMP properties--pdfa-backend {auto,ghostscript,internal} (API:pdfa_backend=). auto, the default, tries OCRmyPDF's own conversion andinternal never uses Ghostscript, so JPEGs pass throughpdfa output types fail with--output-type auto outputs a regular PDF. ghostscript--pdfa-image-compression, --ghostscript-jpeg-quality,--ghostscript-jpeg-maxdpi, and --color-conversion-strategy CMYK,Gray or UseDeviceIndependentColor) now select Ghostscript under autointernal.pikepdf[pdfa] 10.15 or later. The pdfa extrajsonschema, referencing and fonttools, all packaged by--force-ocr now keeps hyperlinks, moving link annotations onto the--deskew and --clean-final. The new mode--mode force-ocr-no-links keeps the old behaviour. {issue}605-dNONATIVEFONTMAP), so output is the same on every1369TZ to choose another61315701024PageInfo.has_visible_text, which excludes invisible text such as anhas_text is unchanged.Fixes
--force-ocr on a scan that already had an invisible OCR layer rasterized961use_threads=False on Windows and macOS) failed on every'OcrOptions' object has no attribute 'tesseract'.1757ocrmypdf.ocr() rejected thresholding method names such astesseract_thresholding='adaptive-otsu'. {issue}1460--rotate-pages --tesseract-timeout 0 detected page orientation but did not778--output-type auto and --force-ocr, output was labelled PDF/A1751605/ToUnicode, which viewers extract through glyph1297dc:contributordc:subject, was lost when Ghostscript made the PDF/A. It is now122016201244948/Lang) was not set for Tesseract's language codesdeu, fra and ces. Languages without a two-letter code now get17490.000-50131235, stopped OCRmyPDF. It now warns and continues, as viewers1054/Rotate stored as a real number, or a cm operator with a non-numeric--force-ocr --ocr-engine none crashed on pages with no images.FileExistsError; on Windows it always did.--output-type auto, the explanation for a file that grew1369'Untitled' (with quotes) for untitled files is1753deskew given inOCR_JSON_SETTINGS when OCR_DESKEW is not set.One column per quarter.
ocrmypdf --version and ocrmypdf --help printed to standard error instead of standard output, so version=$(ocrmypdf --version) came back empty and ocrm
Fixes
ocrmypdf --version and ocrmypdf --help printed to standard errorversion=$(ocrmypdf --version) came backocrmypdf --help | less showed nothing. OCRmyPDF redirects fileOCRmyPDF now requires pikepdf 10.2 or later, up from pikepdf 10. This is the first release that provides pikepdf.NamePath , the Object.as_int() family
Changes
pikepdf.NamePath, the Object.as_int() familyexplicit_conversion(). OCRmyPDFocrmypdf[heic], instead ofpi-heif package we previously depended on ispillow-heif bundles libheif, libde265 and x265ocrmypdf[heic] if you feed HEIC images to1746already_done_ocr) instead of code 2 (input_file),1551stable channel tracks releases again instead of beingFixes
--output-type pdfa file came out as garbagebfchar block, but the CMap specification allows at most1747, fpdf2 issue py-pdf/fpdf2#1952).fi, ff, fl and similar pairs extracted with thecon dentiality, e ects) from --output-type pdfa files1744)./Resources, /Resources /XObject, /Root /AcroForm,/Root /MarkInfo, /Root /Names, /Root /PieceInfo, an annotation's /A,/SMask -- is now tolerated everywhere rather than in the/Subtype no longer raises out of the optimizer.PdfInfo no longer aborts on a /MarkInfo << /Marked 1 >>, where a producer1742)./UserUnit written as a PDF Real is now read exactly rather than via--sidecar now refuses any spelling of the input or output file, not just aocrmypdf in.pdf out.pdf --sidecar ./in.pdf was accepted., .., symlinks and filesystem case. Thanks--deskew and --rotate-pages now warn that they will have no effect when--ocr-engine none, since skew and orientation are measured by1735uv tool install shown inOCR_OUTPUT_DIRECTORY_YEAR_MONTH is now deprecated in favor of OCR_OUTPUT_STRUCTURE=YEAR_MONTH . It is still honored and logs a deprecation warning; if…
Enhancements
--max-ocr-image-mpixels downsamples the image sent to OCR when a page--max-ocr-image-mpixels 8, and recognizes the same text. See the "Memory"watcher.py (the watcher extra) gained a configurable output layout andocrmypdf library API.
OCR_OUTPUT_STRUCTURE setting (--output-structure): FLATYEAR_MONTH ({destination}/{year}/{month}/{filename}, same layout asOCR_OUTPUT_DIRECTORY_YEAR_MONTH=1), or HIERARCHY, whichinput/a/b/c.pdf → output/a/b/c.pdf.OCR_ON_CONFLICT setting (--on-conflict) controls what happensSUFFIX (default)name (1).pdf, name (2).pdf, ... in the OS style; SKIP logsOVERWRITE is the old behavior.SUFFIX, so existing outputOCR_OUTPUT_DIRECTORY and to theOCR_ON_SUCCESS_ARCHIVE; previously the<>:"/\|?* and control_, trailing dots/spaces are stripped,CON, PRN, AUX, NUL, COM1-9,LPT1-9) are prefixed with _.OCR_OUTPUT_DIRECTORY_YEAR_MONTH is now deprecated in favor ofOCR_OUTPUT_STRUCTURE=YEAR_MONTH. It is still honored and logs aOCR_OUTPUT_STRUCTURE wins.Performance
pi-heif already required Pillow 11.1, so the effective floor barely moves).ocrmypdf.ocr(), so a secondocrmypdf.ocr() callsys.modules. Plugins must not rely on being re-executed for each job, and/FlateDecode with a PNG predictor already holds exactlyIDAT chunk holds, so it is now repackaged as a PNG directly--optimize 2). We now decode only on code paths that/Filter) no longer produces a spuriousIndexError internally, which the optimizer's best-effort handler caught andExecutor instances. Executor.pool_lock is retained but no longer acquired,jobs=M may now spawn up tojobsmax_image_mpixels. Pillow's decompression-bomb limit is interpreter-globalFixes
to_pil() lets[WinError 2] The system cannot find the file specified warnings while it searches for Ghostscript and TesseractPROGRAMFILES environment variable pointed to--mode strip failed to remove OCR text layers that OCRmyPDF itself1730). OCRmyPDF grafts its text layer as a Form XObject--mode redo stack a second text layer1732)./Resources inherited from an ancestor/Pages node, rather than only looking at the page's own resources.clean_final on an existing options object in the Python API noclean unset. --clean-final implies --clean, but the rule--jobs 999 now--jobs: Input should be less than or equal to 256. The set ofThe watcher.py watched-folder helper (the watcher extra) has been modernized and security-hardened:
The watcher.py watched-folder helper (the watcher extra) has been
modernized and security-hardened:
watchfiles instead of watchdog. Installingocrmypdf[watcher] now pulls in watchfiles; native OS filesystemOCR_USE_POLLING=1 to forcesys.path,$PATH), ifOCR_JSON_SETTINGS points at a file inside a data directory or one that1715). pikepdf.PasswordError does not derive frompikepdf.PdfError, so it escaped the handler that waits for a file to be1716).See the "Watcher security model" section of the batch processing
documentation for details. Existing
deployments where the data directories are kept separate from the application
are unaffected; deployments that co-located data with the interpreter or its
environment will need to relocate one or the other.
Ghostscript 10.7.0 and later are no longer treated as affected by the JPEG
passthrough truncation bug, which Ghostscript fixed in 10.07.0
({issue}1726). The version check had no upper bound, so users on a fixed
Ghostscript still saw the "JPEG encoding errors" warning and, worse, silently
had every JPEG lossily re-encoded at --optimize 1 (the default) to work
around a bug their Ghostscript did not have. Thanks @zuentec-droid for the
detailed measurements and upstream analysis.
The same JPEG re-encoding workaround no longer applies when Ghostscript did
not produce the file at all. It was previously triggered by the mere presence
of an affected Ghostscript, so --output-type pdf and files converted by the
speculative PDF/A path — neither of which runs Ghostscript — paid the quality
loss for nothing.
Draft release - will be updated when tag is pushed
Draft release - will be updated when tag is pushed
--output-type auto (the default) again produces PDF/A whenever it can, matching OCRmyPDF 16's "PDF/A by default" behavior. It first tries the fast Gho
--output-type auto (the default) again produces PDF/A whenever it can,1561). A consequence is--output-type pdf to skip PDF/A conversion1561). PDF/A requires all fonts to be embedded, so--output-type auto (the default) it produces a--output-type pdfa* it stops with an error rather than emit corrupted--output-type pdf to keep the text layer, or --force-ocr toocrmypdf input.pdf -) is nowprint() calls — ever writing to stdout; a single accidental writeocrmypdf.configure_stdout_protection,ocrmypdf.configure_logging,UnicodeDecodeError when processing a PDF whose/DocumentInfo dictionary contains a /Name key encoded in Latin-1 (or/Saks#e5r. repair_docinfo_nuls now1540). Current pikepdf releases tolerate theseFixed a severe, Windows-specific performance regression in the "Scanning contents" phase, most visible with --redo-ocr ({issue} 1662 ). Since v16.4.3,
--redo-ocr ({issue}1662). Since1361). On Windows, CPython's BufferedReader.read() eagerlyNotoSansArabic[wdth,wght].ttf, the form shipped by Homebrew casks-Regular.ttf/.otf1652).The Docker images now run as a non-root user ( app , uid/gid 1000) by default rather than as root, as a defense-in-depth measure. If you bind-mount a
app, uid/gid 1000) by default--user argument so/data, so files in--workdir.alex-p/tesseract-ocr5 PPA, and the base imagesWhen the optimizer encounters an image it cannot process (for example, an exotic colorspace that cannot be transcoded), it now logs a concise warning
-v 1) ({issue}846).--pdfa-image-compression=auto (the default) now selects lossless image-O0 so Ghostscript no longer transcodes lossless images to-O1 and above, auto continues to defer-O1 (the--pdfa-image-compression=lossless or use -O01124).--pdfa-image-compression=lossless now passes existing JPEG images through/MediaBox, /CropBox, /TrimBox, /ArtBox, /BleedBox) in its1398); rectangles whoseNegativeDimensionError ({issue}1526); and a crop/trim/art/bleed box1400)./Root/PieceInfo/SearchIndex) from its output. This proprietary index,/Thumb image XObject on a page) from its output. OCRmyPDF alters page1688pngmonodpngmono (ordered dithering). It produces-dTextAlphaBits=4 -dGraphicsAlphaBits=4) for the--rasterizer auto)--rasterizer ghostscript or do not have pypdfium2 installed.-v 1), and the --rasterizer help text explains the OCR-quality1439-v 1) so the original wording1566--mode strip, which removes the invisible OCR text layer from a PDF--ocr-engine none --force-ocr, it does not rasterize the1435Added support for the end alias in --pages , denoting the last page of the document. For example, --pages 3-end OCRs from page 3 through the final pag
end alias in --pages, denoting the last page--pages 3-end OCRs from page 3 through1615--ghostscript-jpeg-quality and --ghostscript-jpeg-maxdpi--jpeg-quality remains the recommended file-size control.16851321TesseractConfigError withFileNotFoundError on the missing hOCR output. {issue}1687_exec and subprocess modules toUpdated uv.lock to avoid pinning a vulnerable version of Pillow. {issue} 1666
PIL.Image.MAX_IMAGE_PIXELSmax_image_mpixels. Hostocrmypdf.ocr() now have their setting respected. The CLI16651666Fixed RTL text extraction order in the fpdf2 renderer. Arabic lam-alef ligatures and other multi-character CMap entries were garbled by the bidi algor
1655work_folder not being set in PdfContext options when using1613Added --no-overwrite / -n option to prevent overwriting output files. If the destination file already exists, OCRmyPDF exits with code 5 (OutputFileAc
--no-overwrite / -n option to prevent overwriting output files.
If the destination file already exists, OCRmyPDF exits with code 5
(OutputFileAccessError). {issue}16421635optimize=2 or optimize=3 crash when using the Python API without
explicitly setting jpg_quality or png_quality. {issue}1641verapdf availability check crashing with NotADirectoryError on
some platforms. {issue}1638Fixed Python API ignoring the language parameter, always defaulting to eng. The API now correctly maps language to OcrOptions languages and splits +-s
language parameter, always defaulting to
eng. The API now correctly maps language to OcrOptions languages
and splits +-separated codes (e.g. eng+deu) to match CLI behavior.
{issue}1640tesseract_timeout
defaulted to 0, causing Tesseract to time out immediately. The default is
now None, falling back to the plugin's 180-second timeout. {issue}16361630--image) for the hocrtransform tool,
enabling sandwich PDF output with the fpdf2 renderer. {issue}1634Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →