Download Latest Version v10.16.0 source code.zip (10.2 MB) Google Add to Preferred Sources
Home / v10.14.0
Name Modified Size InfoDownloads / Week
Parent folder
pikepdf-10.14.0-cp314-cp314t-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.sigstore.json 2026-09-26 9.3 kB
pikepdf-10.14.0-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp314-cp314t-musllinux_1_2_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp314-cp314t-musllinux_1_2_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp314-cp314t-win_amd64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp315-cp315t-macosx_15_0_arm64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp315-cp315t-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp315-cp315t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp315-cp315t-musllinux_1_2_aarch64.whl.sigstore.json 2026-09-26 9.3 kB
pikepdf-10.14.0-cp315-cp315t-musllinux_1_2_x86_64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp315-cp315t-win_amd64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0.tar.gz.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp313-cp313-musllinux_1_2_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp313-cp313-musllinux_1_2_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp313-cp313-win_amd64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp314-abi3-macosx_15_0_arm64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp314-abi3-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp314-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp314-abi3-musllinux_1_2_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp314-abi3-musllinux_1_2_x86_64.whl.sigstore.json 2026-09-26 9.3 kB
pikepdf-10.14.0-cp314-abi3-win_amd64.whl.sigstore.json 2026-09-26 9.3 kB
pikepdf-10.14.0-cp314-cp314t-macosx_15_0_arm64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp311-cp311-win_amd64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp312-cp312-macosx_15_0_arm64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp312-cp312-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.sigstore.json 2026-09-26 9.3 kB
pikepdf-10.14.0-cp312-cp312-musllinux_1_2_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp312-cp312-musllinux_1_2_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp312-cp312-win_amd64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp313-cp313-macosx_15_0_arm64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0-cp313-cp313-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp311-cp311-macosx_15_0_arm64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp311-cp311-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.sigstore.json 2026-09-26 9.3 kB
pikepdf-10.14.0-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp311-cp311-musllinux_1_2_aarch64.whl.sigstore.json 2026-09-26 9.2 kB
pikepdf-10.14.0-cp311-cp311-musllinux_1_2_x86_64.whl.sigstore.json 2026-09-26 9.1 kB
pikepdf-10.14.0.tar.gz 2026-09-26 9.9 MB
pikepdf-10.14.0-cp315-cp315t-macosx_15_0_arm64.whl 2026-09-26 2.1 MB
pikepdf-10.14.0-cp315-cp315t-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl 2026-09-26 2.4 MB
pikepdf-10.14.0-cp315-cp315t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-26 2.6 MB
pikepdf-10.14.0-cp315-cp315t-musllinux_1_2_aarch64.whl 2026-09-26 4.0 MB
pikepdf-10.14.0-cp315-cp315t-musllinux_1_2_x86_64.whl 2026-09-26 4.3 MB
pikepdf-10.14.0-cp315-cp315t-win_amd64.whl 2026-09-26 3.9 MB
pikepdf-10.14.0-cp314-abi3-musllinux_1_2_aarch64.whl 2026-09-26 4.0 MB
pikepdf-10.14.0-cp314-abi3-musllinux_1_2_x86_64.whl 2026-09-26 4.3 MB
pikepdf-10.14.0-cp314-abi3-win_amd64.whl 2026-09-26 3.9 MB
pikepdf-10.14.0-cp314-cp314t-macosx_15_0_arm64.whl 2026-09-26 2.1 MB
pikepdf-10.14.0-cp314-cp314t-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl 2026-09-26 2.4 MB
pikepdf-10.14.0-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-26 2.6 MB
pikepdf-10.14.0-cp314-cp314t-musllinux_1_2_aarch64.whl 2026-09-26 4.0 MB
pikepdf-10.14.0-cp314-cp314t-musllinux_1_2_x86_64.whl 2026-09-26 4.3 MB
pikepdf-10.14.0-cp314-cp314t-win_amd64.whl 2026-09-26 3.9 MB
pikepdf-10.14.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-26 2.6 MB
pikepdf-10.14.0-cp313-cp313-musllinux_1_2_aarch64.whl 2026-09-26 4.0 MB
pikepdf-10.14.0-cp313-cp313-musllinux_1_2_x86_64.whl 2026-09-26 4.3 MB
pikepdf-10.14.0-cp313-cp313-win_amd64.whl 2026-09-26 3.8 MB
pikepdf-10.14.0-cp314-abi3-macosx_15_0_arm64.whl 2026-09-26 2.1 MB
pikepdf-10.14.0-cp314-abi3-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl 2026-09-26 2.4 MB
pikepdf-10.14.0-cp314-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-26 2.6 MB
pikepdf-10.14.0-cp312-cp312-macosx_15_0_arm64.whl 2026-09-26 2.1 MB
pikepdf-10.14.0-cp312-cp312-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl 2026-09-26 2.4 MB
pikepdf-10.14.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-26 2.6 MB
pikepdf-10.14.0-cp312-cp312-musllinux_1_2_aarch64.whl 2026-09-26 4.0 MB
pikepdf-10.14.0-cp312-cp312-musllinux_1_2_x86_64.whl 2026-09-26 4.3 MB
pikepdf-10.14.0-cp312-cp312-win_amd64.whl 2026-09-26 3.8 MB
pikepdf-10.14.0-cp313-cp313-macosx_15_0_arm64.whl 2026-09-26 2.1 MB
pikepdf-10.14.0-cp313-cp313-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl 2026-09-26 2.4 MB
pikepdf-10.14.0-cp311-cp311-macosx_15_0_arm64.whl 2026-09-26 2.1 MB
pikepdf-10.14.0-cp311-cp311-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl 2026-09-26 2.4 MB
pikepdf-10.14.0-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-26 2.6 MB
pikepdf-10.14.0-cp311-cp311-musllinux_1_2_aarch64.whl 2026-09-26 4.0 MB
pikepdf-10.14.0-cp311-cp311-musllinux_1_2_x86_64.whl 2026-09-26 4.3 MB
pikepdf-10.14.0-cp311-cp311-win_amd64.whl 2026-09-26 3.8 MB
README.md 2026-09-26 31.8 kB
v10.14.0 source code.tar.gz 2026-09-26 9.9 MB
v10.14.0 source code.zip 2026-09-26 10.1 MB
Totals: 77 Items   145.5 MB 0

Several improvements to explicit conversion mode and NamePath, prompted by the OCRmyPDF project's migration to these APIs (pikepdf.explicit_conversion(), the as_* safe accessors, and NamePath) in a production codebase that reads untrusted, often malformed, PDFs.

Python support

  • Dropped support for Python 3.10, which reaches end of life on October 31,
  • Python 3.11 through 3.14 are supported, and wheels are no longer built for CPython 3.10. pip will keep installing pikepdf 10.13 on Python 3.10.
  • Internal code now uses Python 3.11 features: tomllib, enum.StrEnum, typing.Self, datetime.UTC and the broader ISO 8601 support in datetime.fromisoformat. The compatibility shims for Python 3.10 are gone.
  • CI now tests against CPython 3.15 (release candidate). CPython 3.15 uses the existing cp314-abi3 wheel; free-threaded CPython 3.15 gets its own cp315t wheel, since the stable ABI does not cover free-threaded builds.

Platform support

  • macOS wheels are now built on GitHub's macos-15 runner, since GitHub is deprecating the macos-14 runner. Binary wheels now require macOS 15 or newer.
  • Dropped macOS Intel (x86_64) wheels. We no longer test macOS on Intel, since the platform is end of life. Intel Mac users can build from source, or keep using pikepdf 10.13.

Conversion mode

  • {meth}pikepdf.Pdf.open and {meth}pikepdf.Pdf.new now accept a keyword-only conversion_mode argument ('implicit' or 'explicit'), and the mode can be read or changed afterward through the new {attr}pikepdf.Pdf.conversion_mode property. A Pdf's mode travels with it across threads and does not affect any other document, which makes it the right scope for a library embedded inside a host application. Setting the property to None reverts to inheriting the context-manager/global mode.
  • Added {func}pikepdf.implicit_conversion, a thread-local context manager symmetric with {func}pikepdf.explicit_conversion, so implicit mode can be forced from inside code that runs under explicit mode.
  • Precedence when several scopes are in play: context manager, then per-Pdf mode, then the global default set by {func}pikepdf.set_object_conversion_mode. An object with no owning Pdf (e.g. a bare pikepdf.Dictionary(...) you have not yet attached to a document) resolves through the context-manager/global scopes; once it is inserted into a Pdf it follows that document's mode.
  • Added {meth}pikepdf.Object.get_raw, which behaves like {meth}~pikepdf.Object.get (key, Name, or NamePath, plus a default) but never unboxes the result: it always returns a pikepdf.Object regardless of the current conversion mode. A null stored inside an array comes back as a Null-typed object rather than None; in a dictionary, qpdf treats a key whose value is null as absent, so get_raw returns the default for it, exactly as get does.
  • Added container-level typed getters {meth}~pikepdf.Object.get_int, {meth}~pikepdf.Object.get_bool, {meth}~pikepdf.Object.get_float, {meth}~pikepdf.Object.get_decimal, {meth}~pikepdf.Object.get_dict, and {meth}~pikepdf.Object.get_list, each (key_or_path, default=None). These compose get_raw with the matching as_* accessor, so reading an optional value of uncertain type is a mode-independent one-liner instead of a hand-written with pikepdf.explicit_conversion(): ... wrapper.
  • Added coerce=True to {meth}~pikepdf.Object.as_int, {meth}~pikepdf.Object.as_bool, {meth}~pikepdf.Object.as_float, and {meth}~pikepdf.Object.as_decimal, and to the corresponding get_* getters, for reading values that real-world PDFs encode with a "nearby" type: as_int(coerce=True) accepts a Real (truncated toward zero) and a numeric String; as_bool(coerce=True) accepts a nonzero Integer or Real (fixing the common /Marked 1 case that previously required a fallback to as_int() != 0); as_float/as_decimal with coerce=True accept an Integer and a numeric String, including exponent notation such as "1e-5".
  • pikepdf.Integer and pikepdf.Real now support the ordering comparisons <, <=, > and >= against Python int, float, bool, Decimal, and each other, so box[0] < box[2], sorted(), min() and max() work on explicit-mode values. Real compares by its exact decimal value.
  • Arithmetic on pikepdf.Integer and pikepdf.Real now works between two objects (Real('2.5') + Integer(3)) and with Decimal and bool operands, and ** is supported. Results are the native Python type that implicit mode would have produced: int for Integer with int/Integer, Decimal whenever a Real or Decimal is involved, and float with a float operand. In explicit mode, box[2] - box[0] > 100 therefore works without unboxing, while raising TypeError if the PDF stored a non-number. The result is never a pikepdf object: a number computed from a document is not in the document, and returning an object would let one unmigrated read propagate into code far from the cause.
  • Added {func}pikepdf.unbox, which returns the native Python value of an Integer, Boolean or Real and passes any other value through unchanged. Unlike the as_* accessors it does not require knowing which numeric type a value has, and it accepts values read in either mode, so it is the one-token migration for a read that feeds isinstance, is True, Decimal(), json.dumps or a function documented to return a native type.
  • Added module-level typed conversions {func}pikepdf.as_int, {func}~pikepdf.as_bool, {func}~pikepdf.as_float, {func}~pikepdf.as_decimal, {func}~pikepdf.as_dict, {func}~pikepdf.as_list, {func}~pikepdf.as_str and {func}~pikepdf.as_bytes, each (value, default=None), with keyword-only coerce for the numeric and boolean ones. They are the value-side twins of the get_* typed getters, for an array element, content stream operand or other value already in hand: a pikepdf object converts exactly as its as_* method does, and a native value is accepted only if it is the type implicit mode would have produced for that PDF type, so the result is the same in either conversion mode. pikepdf.as_int(True) gives the default, because a Boolean is not an Integer.
  • Added {meth}~pikepdf.Object.as_str and {meth}~pikepdf.Object.as_bytes, which return a String's decoded text or raw bytes, and raise TypeError (or return a supplied default) for any other type, unlike str() and bytes(), which accept almost anything. Added the matching typed getters {meth}~pikepdf.Object.get_str and {meth}~pikepdf.Object.get_bytes.
  • See {doc}/topics/objects for the full description of scopes and precedence, plus a "Migrating to explicit mode" checklist.
  • Added a new topic page, {doc}/topics/type_safety, explaining why explicit mode exists and setting out the deprecation schedule: v11 warns on implicit conversions, v12 makes explicit mode the default, and v13 removes implicit mode. New code that reads values of uncertain type should prefer the mode-independent getters described above.

NamePath

  • path in obj now tests whether the path can be traversed to a value, and del obj[path] deletes the final component after traversing to its parent, both for name and array-index components.
  • NamePath is now iterable, yielding its individual components (str for names, int for indices) in order, e.g. list(NamePath.A.B[0]) gives ['/A', '/B', 0].
  • NamePath instances now compare and hash by their component sequence, so NamePath.A.B == NamePath.A.B is True and a NamePath can be used as a dict key or stored in a set.
  • isinstance(path, pikepdf.NamePath) now works, and pikepdf.NamePath can be used directly in type annotations (e.g. def f(key: pikepdf.Name | pikepdf.NamePath)) without importing a private name.

Metadata and PDF/A

OCRmyPDF tested its PDF/A output against veraPDF, and found files that pikepdf's metadata handling made invalid.

  • XMP language alternatives such as dc:title and dc:description are now read from their x-default item, as the XMP specification defines, rather than from the first item. Acrobat writes x-default last when a document has several languages, so /Title and /Subject were synced from the wrong language, and PDF/A-1 validators rejected the result.
  • Behavior change: Setting a language alternative now updates its x-default item and keeps the other languages, instead of replacing all of them with the new value. An item that held the same text as the old default is updated too, as Adobe's XMP toolkit does. Assigning several values to a language alternative that pikepdf does not know joins them into one x-default item, instead of writing several x-default items.
  • XMP properties are now read from every top-level rdf:Description, not only those with rdf:about="". Properties in Descriptions with rdf:about="uuid:..." (written by Distiller and older Acrobat) or with no rdf:about could not be read or deleted, and setting one added a duplicate, which made veraPDF reject the whole packet. Setting a property now removes any other occurrences of it, and deleting it removes all of them.
  • Behavior change: When XMP is written, every top-level rdf:Description gets the same rdf:about value, as the XMP specification requires, and empty Descriptions are removed. The value is kept if the non-empty values agree, and made empty if they conflict.
  • Added {attr}pikepdf.models.PdfMetadata.recovered and XmpDocument.recovered, which are True if the XMP was not well-formed and had to be repaired or replaced as it was read.
  • Added {meth}pikepdf.PdfInlineImage.read_raw_bytes, which returns the still-encoded data of an inline image exactly as it appears in the content stream.
  • {meth}pikepdf.Pdf.save now corrects /Count in each node of the page tree to the number of pages beneath it. qpdf repairs a damaged page tree but left /Count as it was, so a wrong /Count in the input was written unchanged, and veraPDF could not validate the file at all.

PDF/A validation and repair

  • Added the {mod}pikepdf.pdfa module, which prepares, checks and saves PDF/A-1b, PDF/A-2b and PDF/A-3b documents. See {ref}pdfa.
  • {func}pikepdf.pdfa.prepare repairs a document in memory: it installs a PDF/A output intent (sRGB by default, or a supplied ICC profile), removes image interpolation and hidden annotations, adds /CIDSet for PDF/A-1, and rewrites the XMP metadata with a PDF/A declaration. It returns a {class}~pikepdf.pdfa.PrepareResult describing what changed: {meth}~pikepdf.pdfa.PrepareResult.describe gives a sentence per change, {meth}~pikepdf.pdfa.PrepareResult.messages the same sentences with a suggested log level, and PrepareResult.xmp_problem why an XMP packet could not be read. XMP properties are kept from every rdf:Description when all share one rdf:about, including the non-empty uuid:... value that Acrobat and pikepdf's own metadata editor write.
  • {func}pikepdf.pdfa.check predicts the verdict for the file that would be written, without writing it, by modelling what qpdf's writer changes (stream filters, trailer, version, encryption and which objects are written).
  • {func}pikepdf.pdfa.save prepares the document, writes it to a temporary file, reopens and validates the bytes written, and only then moves the file into place. If the written file does not pass, it raises {class}~pikepdf.pdfa.PdfaError with the report (whose prepared records the repairs made), and the destination is left untouched.
  • {func}pikepdf.pdfa.resolve_save_kwargs returns the complete {meth}pikepdf.Pdf.save settings for a flavour. Settings that PDF/A forbids or that would change the bytes after the check (such as encryption, normalize_content and fix_metadata_version) are pinned, and a conflicting value raises ValueError.
  • The validator is an allowlist: it approves only constructs it recognizes and knows to conform. A {class}~pikepdf.pdfa.Report has a verdict of 'pass', 'fail' (at least one violation) or 'not_checked' (only constructs the validator does not check, which may be valid PDF/A). save accepts only 'pass'. Each {class}~pikepdf.pdfa.Finding names a rule: a veraPDF rule id such as ISO_19005_2:6.2.8-3, or a pikepdf: id for local policies and unsupported constructs. The validator is not a certification; veraPDF remains the reference.
  • On the 2,906 files of the veraPDF test corpus, converted with prepare and checked against veraPDF 1.30 for each of PDF/A-1b, 2b and 3b, the validator approved no file that veraPDF rejects. For PDF/A-2b it approved 1,363 files, rejected 887 and left 656 not checked; veraPDF accepted 28 of the rejected files (the validator is conservative by design). The prediction from check matched the validation of the written file in every case.
  • The validator and the repairs give the same results whether a document is read in implicit or explicit conversion mode: they work in explicit mode internally, whatever mode the caller uses.
  • {mod}pikepdf.pdfa needs jsonschema, referencing and fontTools, available as the optional extra pip install 'pikepdf[pdfa]'. import pikepdf does not import them; import pikepdf.pdfa raises ImportError naming the extra if they are missing.
  • The PDF/A rule catalogue is derived from the veraPDF validation profiles (CC BY 4.0, veraPDF Consortium); see third-party-licenses/README.md.
  • Added the internal helper pikepdf._io.atomic_write_verified, which writes a file to a temporary location, calls a verification function on it, and moves it into place only if verification succeeds. pikepdf.pdfa.save uses it.

qpdf limits

  • Added {func}pikepdf.settings.get_qpdf_limits and {func}pikepdf.settings.set_qpdf_limits, which read and change qpdf's process-wide limits for hardening against malicious or damaged PDFs: parser nesting depth, error count and container size, the number of stream filters, memory for decoding Flate, DCT, PNG, TIFF and RunLength data, and whether corrupt JPEG data is an error. Previously only some of these limits were reachable, through {meth}pikepdf.JobBuilder.limits. set_qpdf_limits returns the previous values so they can be restored.
  • Added {func}pikepdf.settings.disable_qpdf_default_limits, which lifts qpdf's optional default limits for the rest of the process, and {func}pikepdf.settings.qpdf_limit_errors, which counts how many times any limit has been exceeded.

Behavior changes

  • The minimum required qpdf version is now 12.4.1, and wheels bundle qpdf 12.4.1. Changes inherited from qpdf 12.4:
  • Content stream parsing ({func}pikepdf.parse_content_stream, {meth}pikepdf.Page.parse_contents) now stops after 15 syntax errors, issuing a warning. Instructions after that point are dropped, so parsing and rewriting a badly damaged content stream can lose content. Use pikepdf.settings.set_qpdf_limits(parser_max_errors=0) to parse as much as possible.
  • A page tree nested more than 100 levels deep now raises {class}pikepdf.PdfError when the pages are accessed.
  • /Rotate values outside [0, 360), such as -90, are now normalized when a page is converted to a form XObject, overlaid or rotation-flattened.
  • JSON output of real numbers no longer includes leading zeroes.
  • The progname keyword argument of {class}pikepdf.Job is deprecated and ignored, and passing it issues a DeprecationWarning. It was passed to qpdf's QPDFJob::initializeFromArgv as the name of an environment variable, not the program name, and qpdf 12.4 no longer uses it. The program name is taken from the first item of args, as it always was.
  • Behavior change: Reading PdfImage metadata such as .width, .height or .colorspace whose value has the wrong PDF type (for example a /Width written as a string or name) now raises TypeError, with a message naming the key and the value found, instead of NotImplementedError: Metadata access for Width. The exception is also a NotImplementedError, so existing handlers keep working.
  • Behavior change: A pikepdf.String no longer compares equal to bytes: pikepdf.String('abc') == b'abc' is now False. It still compares equal to the str it decodes to. A String compared equal to both a str and a bytes that are unequal to each other, and so could not hash like both of them; for non-ASCII text it hashed like neither, so {pikepdf.String('héllo'): 1}['héllo'] raised KeyError. A String now hashes like its str. To compare the raw data, use bytes(s) == b'...' or {meth}~pikepdf.Object.as_bytes.
  • Behavior change: A pikepdf.Name no longer compares equal to bytes: pikepdf.Name('/Foo') == b'/Foo' is now False. It still compares equal to the str of its UTF-8 bytes, so obj.Type == '/Page' works as before, and it now hashes like that str, so {pikepdf.Name('/héllo'): 1}['/héllo'] no longer raises KeyError. As PDF 2.0 (ISO 32000-2, 7.3.5) requires, names are compared byte for byte, with no Unicode normalization. Use bytes(name) == b'...' to compare the raw bytes.
  • Behavior change: A pikepdf.Real combined with an int, Integer, or Decimal, or negated with unary -/+/abs(), now yields a Decimal (previously a float, or TypeError for int operands other than /). A Real with a float operand still yields a float. Division of an Integer or Real by zero now raises ZeroDivisionError, as for Python numbers, instead of ValueError.
  • Behavior change: {class}pikepdf.Matrix now raises TypeError, not ValueError, for a pikepdf.Object that is not a matrix or an ObjectList with a non-numeric element, matching what a native argument of the wrong type raises, so an except TypeError around Matrix(*operands) behaves the same in either conversion mode. Size errors (must have 6 elements) remain ValueError.
  • Behavior change: bool() on a pikepdf.Integer or pikepdf.Real is now by value (bool(pikepdf.Integer(0)) is False), instead of raising NotImplementedError: code is unreachable.
  • Behavior change: {meth}~pikepdf.Object.as_dict and {meth}~pikepdf.Object.as_list now accept a default argument and raise TypeError on a type mismatch, instead of raising pikepdf.PdfError with no way to supply a default.
  • Behavior change: Objects created in Python are now adopted by the Pdf they are inserted into, so pdf.Root.X = 42 followed by pdf.Root.get_raw('/X').is_owned_by(pdf) is True, matching objects parsed from a file. Scalars (Name, String, Integer, ...) are adopted as a copy, so a handle held in a constant can still be inserted into several documents. A Dictionary or Array is adopted in place and keeps its aliasing with the document; inserting the same container into a second Pdf now raises ForeignObjectError, as it already did for containers parsed from a file. Construct a new object, or use {meth}pikepdf.Pdf.copy_foreign. Adoption is not permanent: removing the object from the document -- deleting or replacing the dictionary key or array element that held it, or destroying the document -- releases the claim, so the object can then be inserted into another Pdf. If the same object was reachable under two keys, removing either one releases it; its value is unaffected and still readable through the remaining reference.
  • Behavior change: Objects copied into a document are adopted by it too. The direct children of an object produced by {meth}pikepdf.Pdf.copy_foreign, {meth}pikepdf.Object.with_same_owner_as, {meth}pikepdf.Pdf.make_indirect, or by appending/inserting a page into Pdf.pages, and values inserted through NameTree/NumberTree __setitem__, now report the destination document as their owner and follow its conversion_mode.
  • Behavior change: {meth}pikepdf.Object.with_same_owner_as and {meth}pikepdf.Pdf.make_indirect now raise ForeignObjectError for a direct object that already belongs to a different Pdf, instead of silently retagging an object that the other document still references. Remove it from that document first, or make it indirect there and use {meth}pikepdf.Pdf.copy_foreign.
  • Behavior change: {meth}~pikepdf.Object.as_int with coerce=True now returns a supplied default for a value that does not fit in a 64-bit PDF integer, instead of raising OverflowError. Called without a default it still raises OverflowError.
  • Behavior change: repr() of a pikepdf.Object now honors the effective conversion mode of the object being displayed (context manager, then its owning Pdf, then the global setting), rather than only the global setting.
  • Behavior change: {func}pikepdf.set_object_conversion_mode now raises ValueError for a value other than 'implicit' or 'explicit', instead of silently accepting it.

Performance

  • {meth}pikepdf.Pdf.save writes directly to the file descriptor when saving to a filename or to a plain binary file object from {func}open, instead of calling the stream's write() method for every chunk of output. Saving a large document is about twice as fast. Other streams, such as BytesIO, pipes and subclasses of the io classes, are written through write() as before.
  • isinstance() checks against {class}pikepdf.Dictionary, {class}pikepdf.Stream, {class}pikepdf.Integer and the other object classes are about five times faster.
  • {func}pikepdf.unbox is now implemented in C++ and is about 40 times faster. {func}pikepdf.as_int, {func}pikepdf.as_float and {func}pikepdf.as_decimal answer for a native int or Decimal without constructing a PDF object, and converting a Real to a Decimal no longer imports the decimal module each time.

Fixes

  • {func}pikepdf.models.metadata.decode_pdf_date now accepts every PDF date form the specification allows, in which all fields after the year are optional. D:202001011230 (no seconds) was misread as 12:03, and D:2020010112 (no minutes) and offsets in hours only, such as -08', raised ValueError, so {attr}pikepdf.Pdf.docinfo dates in these forms were not synchronized to XMP.
  • {meth}pikepdf.Pdf.save now raises ValueError when the destination is the stream the Pdf was opened from, or another stream on the same file, as it already did for the input file's path. Previously the new file was written into the input at the stream's current position, which produced a corrupt file and could corrupt the open Pdf, since qpdf reads its input lazily. Open with allow_overwriting_input=True to save over the input.
  • Setting an XMP property that a file stored as plain text, such as <dc:title>Old</dc:title>, now replaces the text. Previously the new value was added alongside the old text, which was still returned when the property was read back.
  • {attr}pikepdf.PdfImage.filter_decodeparms and {attr}pikepdf.PdfImage.decode_parms no longer raise AttributeError when an image's /DecodeParms is a bare number or boolean, in either conversion mode. A /DecodeParms that is neither a dictionary nor an array is now ignored, as qpdf does, so each filter gets empty parameters.
  • An image /Decode that is not an array (a bare number, boolean, name or dictionary) is now ignored in favour of the default decode array, instead of raising AttributeError when the image is read or converted.
  • {meth}pikepdf.Stream.write now raises TypeError rather than AttributeError when filter or decode_parms is a native Python value such as an int, str or dict.
  • {meth}pikepdf.Pdf.save to an existing file that is not a regular file, such as /dev/null, a FIFO or a character device, now writes into it directly. Previously the temporary file was renamed over it, so when run as root, saving to /dev/null replaced the device with a regular file.
  • {meth}pikepdf.Matrix.inverse now returns the correct translation for matrices whose determinant is not 1. Previously the e and f terms were not divided by the determinant, so the result was not a true inverse for any matrix that both scaled and translated. Matrices built only from rotations and translations were unaffected. (#743)
  • Storing a Python int outside the signed 64-bit range of a PDF integer, by assignment or in Array(...), Dictionary(...) or Integer(...), now raises OverflowError instead of RuntimeError: std::bad_cast (or a misleading TypeError).
  • PdfImage, PdfImageBase, PdfJpxImage, PdfInlineImage and PaletteData once again report their module as pikepdf.models.image, the path they are imported from, instead of the private pikepdf.models.image._classes introduced when that module became a package in v10.10.
  • The examples for {func}pikepdf.set_object_conversion_mode and {func}pikepdf.explicit_conversion no longer open a nonexistent test.pdf, and the global-mode example restores the default mode when it is done.
  • str() of a pikepdf.Integer, Boolean or Real now gives the value -- '42', 'True', '1.50' -- rather than the object's repr. In implicit mode a scalar arrives as int/bool/Decimal and str() never reached the object, so every f-string, log line and string concatenation in code reading PDF values silently changed meaning under explicit mode.
  • hash() now works on pikepdf.Integer, Boolean and Real (and on a null object), instead of raising RuntimeError: don't know how to hash this. A scalar hashes like the Python value it compares equal to, so Real('1.0'), Decimal('1.0') and 1 agree, and a scalar can be a dict key or a set member as it could in implicit mode.
  • {meth}pikepdf.Pdf.make_indirect and {meth}pikepdf.Object.with_same_owner_as no longer turn a direct scalar (Name, String, Integer, Real, Boolean, Operator) that you pass them into an indirect object in place. They make a copy indirect and return it, leaving your object direct and unowned. Previously the object you held silently became an indirect object of that Pdf, after which hash() raised and any dict or set already holding it broke; a module-level constant such as FOO = Name.Foo became tied to one document. Arrays, dictionaries and streams are still made indirect in place, so the object you passed stays aliased with the one in the document.
  • In explicit mode, repr() of an Array, Dictionary or Stream now names the scalars nested inside it -- pikepdf.Real('42.42') rather than '42.42', which was indistinguishable from a PDF string and did not survive eval(repr(obj)).
  • pikepdf's own higher-level APIs now work under explicit conversion mode. They read values out of PDF objects and compute with them, and every such read assumed the implicit-mode native type. Fixed in image extraction (PdfImage raised NotImplementedError for any indexed or /DeviceN colorspace), page labels (page.label restarted numbering at 1 and warned about a valid /St), outlines (a closed item read as open, flags came back empty, and saving wrote the wrong /Count), form appearance generation (TypeError laying out multiline and combed text fields), get_objects_with_ctm (a malformed cm operator raised instead of being skipped), SimpleFont metrics, and Action.new_window.
  • PdfImage metadata now reads the same in either conversion mode. A real /Width or /Height or a boolean /BitsPerComponent raised NotImplementedError in explicit mode, and a real inside a /ColorSpace array raised it in implicit mode.
  • SimpleFont.ascent, descent and unscaled_char_width() now return a Decimal as documented, in either conversion mode. Font metrics stored as PDF integers were previously returned as int.
  • Iterating a NamePath (e.g. list(path)) no longer falls back to the legacy __getitem__(0), (1), ... protocol, which never raised IndexError and so iterated forever, exhausting memory.
  • NamePath in obj no longer silently answers False for every path; see above.
  • Fixed a stale NamePath['/A']['/B'].C[0] example in the type stub, which did not type-check against the runtime behavior (an instance's __getitem__ only accepts an int). The correct form, matching the compiled docstring and the {doc}/topics/namepath documentation, is NamePath['/A']('/B').C[0].
  • A direct object parsed from a file, then removed from the document (for example box = page.obj.get_raw('/MediaBox') followed by del page.obj['/MediaBox']) and used after the document was closed, kept a dangling reference to the freed document and could crash the interpreter on repr() or when inserted elsewhere. Removing an object from a document now releases its association with it, and a Pdf disconnects every object it adopted as it is destroyed.

Internals

  • Linux wheel jobs now cache the compiled qpdf per build image with actions/cache, so jobs that share an image no longer each download and compile the same qpdf. See the build process notes.
  • Fixed the Linux wheel build script's AlmaLinux detection, which was a malformed shell test that always evaluated false.
  • The type stubs for the C++ extension module, until now a single 4,300-line src/pikepdf/_core.pyi, are a stub-only package src/pikepdf/_core/ split into one stub per translation unit: _core/_matrix.pyi covers src/core/matrix.cpp, _core/_page.pyi covers src/core/page.cpp, and so on, with _core/__init__.pyi re-exporting the lot and documenting the layout. Nothing changes at runtime or in the public API -- pikepdf._core is still one extension module and import pikepdf._core still resolves to it. Type checkers now name a type by its defining stub in messages (pikepdf._core._object.Object rather than pikepdf._core.Object); pikepdf.Object remains the name to write in annotations.
  • API documentation for the C++ extension now lives only in the src/pikepdf/_core/ stubs, which are what Sphinx, type checkers and IDEs read. Docstrings attached to the C++ bindings were never published and had drifted from the stubs; they are merged into the stubs and removed from C++, and a test now fails if a binding gains a docstring or a public C++ name is missing from the stubs. help() on a C++ method now shows only its signature, as it already did for most of them.

Documentation

  • Restored the full {class}pikepdf.Stream and {class}pikepdf.Dictionary documentation, including constructor arguments and examples, which had been reduced to a single line when those classes moved to C++.

  • Clarified guidance on Pdf.save(..., deterministic_id=) and static_id=, and the matching JobBuilder.deterministic_id() and JobBuilder.static_id(). deterministic_id gives reproducible, production-safe /ID values; static_id sets the same dummy /ID in every file and is for testing only. static_id now appears last in the Pdf.save() signature. All save() options are keyword-only, so existing code is unaffected.

Source: README.md, updated 2026-09-26