| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| ocrmypdf-17.13.0-py3-none-any.whl.sigstore.json | 2026-09-28 | 11.2 kB | |
| ocrmypdf-17.13.0.tar.gz.sigstore.json | 2026-09-28 | 11.2 kB | |
| ocrmypdf-17.13.0-py3-none-any.whl | 2026-09-28 | 561.9 kB | |
| ocrmypdf-17.13.0.tar.gz | 2026-09-28 | 7.7 MB | |
| README.md | 2026-09-28 | 6.5 kB | |
| v17.13.0 source code.tar.gz | 2026-09-28 | 7.7 MB | |
| v17.13.0 source code.zip | 2026-09-28 | 8.0 MB | |
| Totals: 7 Items | 24.0 MB | 92 | |
Changes
- PDF/A made without Ghostscript ("speculative" conversion) is now validated
with pikepdf's PDF/A support (
pikepdf.pdfa) instead of veraPDF, and the final output is validated again after the metadata and optimization steps. veraPDF is no longer used, so the Ghostscript-free path is available whether or not veraPDF is installed, including in the Docker image. Files the validator does not approve go to Ghostscript as before; run with-v1to see why. - Speculative conversion now supports PDF/A-1b, and repairs common problems
that previously always sent a file to Ghostscript: it replaces other output
intents with sRGB, removes image interpolation flags, sets the Print flag on
annotations, adds the
/CIDSetPDF/A-1 requires, and removes XMP properties and hidden annotations that PDF/A does not permit. - New option
--pdfa-backend {auto,ghostscript,internal}(API:pdfa_backend=).auto, the default, tries OCRmyPDF's own conversion and then Ghostscript.internalnever uses Ghostscript, so JPEGs pass through unchanged; if the result is not approved, thepdfaoutput types fail with exit code 10 and--output-type autooutputs a regular PDF.ghostscriptalways uses Ghostscript. Options that only Ghostscript implements (--pdfa-image-compression,--ghostscript-jpeg-quality,--ghostscript-jpeg-maxdpi, and--color-conversion-strategyCMYK,GrayorUseDeviceIndependentColor) now select Ghostscript underautoand are an error withinternal. - OCRmyPDF now requires
pikepdf[pdfa]10.15 or later. Thepdfaextra brings injsonschema,referencingandfonttools, all packaged by Debian and Red Hat. --force-ocrnow keeps hyperlinks, moving link annotations onto the rasterized page; so do--deskewand--clean-final. The new mode--mode force-ocr-no-linkskeeps the old behaviour. {issue}605- When Ghostscript substitutes fonts that are not embedded, it now always uses
its own fonts (
-dNONATIVEFONTMAP), so output is the same on every platform; on macOS, native font lookup could embed tens of megabytes of system fonts. OCRmyPDF warns when fonts other than the standard 14 are substituted, or when a substitute loses bold or italic styling. {issue}1369 - A creation date without a time zone is now taken to be local time, and the
output's dates get that zone, with a warning. Set
TZto choose another zone. - In the default mode, a page that already has text now stops OCRmyPDF
straight after the initial scan, instead of after OCRing the pages before
it. The error names the page. {issue}
613 - The check of the output file is much faster: each stream is checked in the
cheapest way for its compression instead of being fully decoded, and a
progress bar is shown. JPEG 2000 and CCITT images are now checked too, but
JPEGs that are corrupt without being truncated are no longer detected.
{issue}
1570 - OCRmyPDF now ships a PyInstaller hook, so apps that use it can be bundled
without extra options. Tesseract and Ghostscript must still be installed.
{issue}
1024 - New
PageInfo.has_visible_text, which excludes invisible text such as an OCR layer.has_textis unchanged.
Fixes
--force-ocron a scan that already had an invisible OCR layer rasterized it at no less than 400 dpi, as if the text were visible, which could multiply the file size. {issue}961- Process workers (
use_threads=Falseon Windows and macOS) failed on every page with'OcrOptions' object has no attribute 'tesseract'. {issue}1757 ocrmypdf.ocr()rejected thresholding method names such astesseract_thresholding='adaptive-otsu'. {issue}1460--rotate-pages --tesseract-timeout 0detected page orientation but did not rotate the pages. The cookbook now recommends this combination for rotating or deskewing without OCR. {issue}778- With
--output-type autoand--force-ocr, output was labelled PDF/A without being validated or given the PDF/A declarations when veraPDF was not installed, as in the Docker image. {issue}1751 - Ghostscript's PDF/A conversion deleted most hyperlinks (those without the
Print flag). Hidden annotations, which PDF/A does not permit, are now
removed before Ghostscript runs, with the same warning as speculative
conversion, instead of silently. {issue}
605 - Text in fonts without
/ToUnicode, which viewers extract through glyph names, was garbled or lost when Ghostscript made the PDF/A. {issue}1297 - XMP metadata with no document-info equivalent, such as
dc:contributoranddc:subject, was lost when Ghostscript made the PDF/A. It is now copied with its whole value, including every language of a multilingual property. {issue}1220 - Images in PDFs from iText and pdftk were recompressed without their Flate
predictor, growing the output by about 30%. {issue}
1620 - Copied text from vertical Japanese and other vertical or rotated OCR lines
had spurious spaces and characters out of order. {issue}
1244 - Text set in the glyphless fallback font, used when no installed font covers a script (e.g. CJK), lost characters when selected or copied in Chrome and other pdfium-based viewers.
- A page made of full-page images at different resolutions was rasterized at
a weighted average resolution rather than the highest. {issue}
948 - The document language (
/Lang) was not set for Tesseract's language codes such asdeu,fraandces. Languages without a two-letter code now get their three-letter code. {issue}1749 - Some LaTeX PDFs with malformed numbers in a content stream, such as
0.000-50131235, stopped OCRmyPDF. It now warns and continues, as viewers do. {issue}1054 - A
/Rotatestored as a real number, or acmoperator with a non-numeric operand, crashed the page scan. --force-ocr --ocr-engine nonecrashed on pages with no images.- When OCRmyPDF's own PDF/A failed its final check, the Ghostscript fallback
could fail with
FileExistsError; on Windows it always did. - With the default
--output-type auto, the explanation for a file that grew now points to PDF/A conversion. {issue}1369 - Ghostscript 10.08's title
'Untitled'(with quotes) for untitled files is removed again. - The Docker image's web service could not find its Streamlit script. The
Docker documentation for the web service is updated. {issue}
1753 - The Docker watcher no longer turns off
deskewgiven inOCR_JSON_SETTINGSwhenOCR_DESKEWis not set.