Several improvements to explicit conversion mode and NamePath, prompted by
the OCRmyPDF project's migration to these APIs
(pikepdf.explicit_conversion(), the as_* safe accessors, and NamePath)
in a production codebase that reads untrusted, often malformed, PDFs.
Python support
- Dropped support for Python 3.10, which reaches end of life on October 31,
- Python 3.11 through 3.14 are supported, and wheels are no longer built for CPython 3.10. pip will keep installing pikepdf 10.13 on Python 3.10.
- Internal code now uses Python 3.11 features:
tomllib,enum.StrEnum,typing.Self,datetime.UTCand the broader ISO 8601 support indatetime.fromisoformat. The compatibility shims for Python 3.10 are gone. - CI now tests against CPython 3.15 (release candidate). CPython 3.15 uses the
existing
cp314-abi3wheel; free-threaded CPython 3.15 gets its owncp315twheel, since the stable ABI does not cover free-threaded builds.
Platform support
- macOS wheels are now built on GitHub's macos-15 runner, since GitHub is deprecating the macos-14 runner. Binary wheels now require macOS 15 or newer.
- Dropped macOS Intel (x86_64) wheels. We no longer test macOS on Intel, since the platform is end of life. Intel Mac users can build from source, or keep using pikepdf 10.13.
Conversion mode
- {meth}
pikepdf.Pdf.openand {meth}pikepdf.Pdf.newnow accept a keyword-onlyconversion_modeargument ('implicit'or'explicit'), and the mode can be read or changed afterward through the new {attr}pikepdf.Pdf.conversion_modeproperty. APdf's mode travels with it across threads and does not affect any other document, which makes it the right scope for a library embedded inside a host application. Setting the property toNonereverts to inheriting the context-manager/global mode. - Added {func}
pikepdf.implicit_conversion, a thread-local context manager symmetric with {func}pikepdf.explicit_conversion, so implicit mode can be forced from inside code that runs under explicit mode. - Precedence when several scopes are in play: context manager, then per-
Pdfmode, then the global default set by {func}pikepdf.set_object_conversion_mode. An object with no owningPdf(e.g. a barepikepdf.Dictionary(...)you have not yet attached to a document) resolves through the context-manager/global scopes; once it is inserted into aPdfit follows that document's mode. - Added {meth}
pikepdf.Object.get_raw, which behaves like {meth}~pikepdf.Object.get(key,Name, orNamePath, plus adefault) but never unboxes the result: it always returns apikepdf.Objectregardless of the current conversion mode. A null stored inside an array comes back as aNull-typed object rather thanNone; in a dictionary, qpdf treats a key whose value is null as absent, soget_rawreturns the default for it, exactly asgetdoes. - Added container-level typed getters {meth}
~pikepdf.Object.get_int, {meth}~pikepdf.Object.get_bool, {meth}~pikepdf.Object.get_float, {meth}~pikepdf.Object.get_decimal, {meth}~pikepdf.Object.get_dict, and {meth}~pikepdf.Object.get_list, each(key_or_path, default=None). These composeget_rawwith the matchingas_*accessor, so reading an optional value of uncertain type is a mode-independent one-liner instead of a hand-writtenwith pikepdf.explicit_conversion(): ...wrapper. - Added
coerce=Trueto {meth}~pikepdf.Object.as_int, {meth}~pikepdf.Object.as_bool, {meth}~pikepdf.Object.as_float, and {meth}~pikepdf.Object.as_decimal, and to the correspondingget_*getters, for reading values that real-world PDFs encode with a "nearby" type:as_int(coerce=True)accepts aReal(truncated toward zero) and a numericString;as_bool(coerce=True)accepts a nonzeroIntegerorReal(fixing the common/Marked 1case that previously required a fallback toas_int() != 0);as_float/as_decimalwithcoerce=Trueaccept anIntegerand a numericString, including exponent notation such as"1e-5". pikepdf.Integerandpikepdf.Realnow support the ordering comparisons<,<=,>and>=against Pythonint,float,bool,Decimal, and each other, sobox[0] < box[2],sorted(),min()andmax()work on explicit-mode values.Realcompares by its exact decimal value.- Arithmetic on
pikepdf.Integerandpikepdf.Realnow works between two objects (Real('2.5') + Integer(3)) and withDecimalandbooloperands, and**is supported. Results are the native Python type that implicit mode would have produced:intforIntegerwithint/Integer,Decimalwhenever aRealorDecimalis involved, andfloatwith afloatoperand. In explicit mode,box[2] - box[0] > 100therefore works without unboxing, while raisingTypeErrorif the PDF stored a non-number. The result is never a pikepdf object: a number computed from a document is not in the document, and returning an object would let one unmigrated read propagate into code far from the cause. - Added {func}
pikepdf.unbox, which returns the native Python value of anInteger,BooleanorRealand passes any other value through unchanged. Unlike theas_*accessors it does not require knowing which numeric type a value has, and it accepts values read in either mode, so it is the one-token migration for a read that feedsisinstance,is True,Decimal(),json.dumpsor a function documented to return a native type. - Added module-level typed conversions {func}
pikepdf.as_int, {func}~pikepdf.as_bool, {func}~pikepdf.as_float, {func}~pikepdf.as_decimal, {func}~pikepdf.as_dict, {func}~pikepdf.as_list, {func}~pikepdf.as_strand {func}~pikepdf.as_bytes, each(value, default=None), with keyword-onlycoercefor the numeric and boolean ones. They are the value-side twins of theget_*typed getters, for an array element, content stream operand or other value already in hand: a pikepdf object converts exactly as itsas_*method does, and a native value is accepted only if it is the type implicit mode would have produced for that PDF type, so the result is the same in either conversion mode.pikepdf.as_int(True)gives the default, because a Boolean is not an Integer. - Added {meth}
~pikepdf.Object.as_strand {meth}~pikepdf.Object.as_bytes, which return aString's decoded text or raw bytes, and raiseTypeError(or return a supplied default) for any other type, unlikestr()andbytes(), which accept almost anything. Added the matching typed getters {meth}~pikepdf.Object.get_strand {meth}~pikepdf.Object.get_bytes. - See {doc}
/topics/objectsfor the full description of scopes and precedence, plus a "Migrating to explicit mode" checklist. - Added a new topic page, {doc}
/topics/type_safety, explaining why explicit mode exists and setting out the deprecation schedule: v11 warns on implicit conversions, v12 makes explicit mode the default, and v13 removes implicit mode. New code that reads values of uncertain type should prefer the mode-independent getters described above.
NamePath
path in objnow tests whether the path can be traversed to a value, anddel obj[path]deletes the final component after traversing to its parent, both for name and array-index components.NamePathis now iterable, yielding its individual components (strfor names,intfor indices) in order, e.g.list(NamePath.A.B[0])gives['/A', '/B', 0].NamePathinstances now compare and hash by their component sequence, soNamePath.A.B == NamePath.A.BisTrueand aNamePathcan be used as a dict key or stored in aset.isinstance(path, pikepdf.NamePath)now works, andpikepdf.NamePathcan be used directly in type annotations (e.g.def f(key: pikepdf.Name | pikepdf.NamePath)) without importing a private name.
Metadata and PDF/A
OCRmyPDF tested its PDF/A output against veraPDF, and found files that pikepdf's metadata handling made invalid.
- XMP language alternatives such as
dc:titleanddc:descriptionare now read from theirx-defaultitem, as the XMP specification defines, rather than from the first item. Acrobat writesx-defaultlast when a document has several languages, so/Titleand/Subjectwere synced from the wrong language, and PDF/A-1 validators rejected the result. - Behavior change: Setting a language alternative now updates its
x-defaultitem and keeps the other languages, instead of replacing all of them with the new value. An item that held the same text as the old default is updated too, as Adobe's XMP toolkit does. Assigning several values to a language alternative that pikepdf does not know joins them into onex-defaultitem, instead of writing severalx-defaultitems. - XMP properties are now read from every top-level
rdf:Description, not only those withrdf:about="". Properties in Descriptions withrdf:about="uuid:..."(written by Distiller and older Acrobat) or with nordf:aboutcould not be read or deleted, and setting one added a duplicate, which made veraPDF reject the whole packet. Setting a property now removes any other occurrences of it, and deleting it removes all of them. - Behavior change: When XMP is written, every top-level
rdf:Descriptiongets the samerdf:aboutvalue, as the XMP specification requires, and empty Descriptions are removed. The value is kept if the non-empty values agree, and made empty if they conflict. - Added {attr}
pikepdf.models.PdfMetadata.recoveredandXmpDocument.recovered, which areTrueif the XMP was not well-formed and had to be repaired or replaced as it was read. - Added {meth}
pikepdf.PdfInlineImage.read_raw_bytes, which returns the still-encoded data of an inline image exactly as it appears in the content stream. - {meth}
pikepdf.Pdf.savenow corrects/Countin each node of the page tree to the number of pages beneath it. qpdf repairs a damaged page tree but left/Countas it was, so a wrong/Countin the input was written unchanged, and veraPDF could not validate the file at all.
PDF/A validation and repair
- Added the {mod}
pikepdf.pdfamodule, which prepares, checks and saves PDF/A-1b, PDF/A-2b and PDF/A-3b documents. See {ref}pdfa. - {func}
pikepdf.pdfa.preparerepairs a document in memory: it installs a PDF/A output intent (sRGB by default, or a supplied ICC profile), removes image interpolation and hidden annotations, adds/CIDSetfor PDF/A-1, and rewrites the XMP metadata with a PDF/A declaration. It returns a {class}~pikepdf.pdfa.PrepareResultdescribing what changed: {meth}~pikepdf.pdfa.PrepareResult.describegives a sentence per change, {meth}~pikepdf.pdfa.PrepareResult.messagesthe same sentences with a suggested log level, andPrepareResult.xmp_problemwhy an XMP packet could not be read. XMP properties are kept from everyrdf:Descriptionwhen all share onerdf:about, including the non-emptyuuid:...value that Acrobat and pikepdf's own metadata editor write. - {func}
pikepdf.pdfa.checkpredicts the verdict for the file that would be written, without writing it, by modelling what qpdf's writer changes (stream filters, trailer, version, encryption and which objects are written). - {func}
pikepdf.pdfa.saveprepares the document, writes it to a temporary file, reopens and validates the bytes written, and only then moves the file into place. If the written file does not pass, it raises {class}~pikepdf.pdfa.PdfaErrorwith the report (whosepreparedrecords the repairs made), and the destination is left untouched. - {func}
pikepdf.pdfa.resolve_save_kwargsreturns the complete {meth}pikepdf.Pdf.savesettings for a flavour. Settings that PDF/A forbids or that would change the bytes after the check (such asencryption,normalize_contentandfix_metadata_version) are pinned, and a conflicting value raisesValueError. - The validator is an allowlist: it approves only constructs it recognizes and
knows to conform. A {class}
~pikepdf.pdfa.Reporthas a verdict of'pass','fail'(at least one violation) or'not_checked'(only constructs the validator does not check, which may be valid PDF/A).saveaccepts only'pass'. Each {class}~pikepdf.pdfa.Findingnames a rule: a veraPDF rule id such asISO_19005_2:6.2.8-3, or apikepdf:id for local policies and unsupported constructs. The validator is not a certification; veraPDF remains the reference. - On the 2,906 files of the veraPDF test corpus, converted with
prepareand checked against veraPDF 1.30 for each of PDF/A-1b, 2b and 3b, the validator approved no file that veraPDF rejects. For PDF/A-2b it approved 1,363 files, rejected 887 and left 656 not checked; veraPDF accepted 28 of the rejected files (the validator is conservative by design). The prediction fromcheckmatched the validation of the written file in every case. - The validator and the repairs give the same results whether a document is read in implicit or explicit conversion mode: they work in explicit mode internally, whatever mode the caller uses.
- {mod}
pikepdf.pdfaneeds jsonschema, referencing and fontTools, available as the optional extrapip install 'pikepdf[pdfa]'.import pikepdfdoes not import them;import pikepdf.pdfaraisesImportErrornaming the extra if they are missing. - The PDF/A rule catalogue is derived from the veraPDF validation
profiles (CC BY 4.0, veraPDF Consortium); see
third-party-licenses/README.md. - Added the internal helper
pikepdf._io.atomic_write_verified, which writes a file to a temporary location, calls a verification function on it, and moves it into place only if verification succeeds.pikepdf.pdfa.saveuses it.
qpdf limits
- Added {func}
pikepdf.settings.get_qpdf_limitsand {func}pikepdf.settings.set_qpdf_limits, which read and change qpdf's process-wide limits for hardening against malicious or damaged PDFs: parser nesting depth, error count and container size, the number of stream filters, memory for decoding Flate, DCT, PNG, TIFF and RunLength data, and whether corrupt JPEG data is an error. Previously only some of these limits were reachable, through {meth}pikepdf.JobBuilder.limits.set_qpdf_limitsreturns the previous values so they can be restored. - Added {func}
pikepdf.settings.disable_qpdf_default_limits, which lifts qpdf's optional default limits for the rest of the process, and {func}pikepdf.settings.qpdf_limit_errors, which counts how many times any limit has been exceeded.
Behavior changes
- The minimum required qpdf version is now 12.4.1, and wheels bundle qpdf 12.4.1. Changes inherited from qpdf 12.4:
- Content stream parsing ({func}
pikepdf.parse_content_stream, {meth}pikepdf.Page.parse_contents) now stops after 15 syntax errors, issuing a warning. Instructions after that point are dropped, so parsing and rewriting a badly damaged content stream can lose content. Usepikepdf.settings.set_qpdf_limits(parser_max_errors=0)to parse as much as possible. - A page tree nested more than 100 levels deep now raises
{class}
pikepdf.PdfErrorwhen the pages are accessed. /Rotatevalues outside[0, 360), such as-90, are now normalized when a page is converted to a form XObject, overlaid or rotation-flattened.- JSON output of real numbers no longer includes leading zeroes.
- The
prognamekeyword argument of {class}pikepdf.Jobis deprecated and ignored, and passing it issues aDeprecationWarning. It was passed to qpdf'sQPDFJob::initializeFromArgvas the name of an environment variable, not the program name, and qpdf 12.4 no longer uses it. The program name is taken from the first item ofargs, as it always was. - Behavior change: Reading
PdfImagemetadata such as.width,.heightor.colorspacewhose value has the wrong PDF type (for example a/Widthwritten as a string or name) now raisesTypeError, with a message naming the key and the value found, instead ofNotImplementedError: Metadata access for Width. The exception is also aNotImplementedError, so existing handlers keep working. - Behavior change: A
pikepdf.Stringno longer compares equal tobytes:pikepdf.String('abc') == b'abc'is nowFalse. It still compares equal to thestrit decodes to. AStringcompared equal to both astrand abytesthat are unequal to each other, and so could not hash like both of them; for non-ASCII text it hashed like neither, so{pikepdf.String('héllo'): 1}['héllo']raisedKeyError. AStringnow hashes like itsstr. To compare the raw data, usebytes(s) == b'...'or {meth}~pikepdf.Object.as_bytes. - Behavior change: A
pikepdf.Nameno longer compares equal tobytes:pikepdf.Name('/Foo') == b'/Foo'is nowFalse. It still compares equal to thestrof its UTF-8 bytes, soobj.Type == '/Page'works as before, and it now hashes like thatstr, so{pikepdf.Name('/héllo'): 1}['/héllo']no longer raisesKeyError. As PDF 2.0 (ISO 32000-2, 7.3.5) requires, names are compared byte for byte, with no Unicode normalization. Usebytes(name) == b'...'to compare the raw bytes. - Behavior change: A
pikepdf.Realcombined with anint,Integer, orDecimal, or negated with unary-/+/abs(), now yields aDecimal(previously afloat, orTypeErrorforintoperands other than/). ARealwith afloatoperand still yields afloat. Division of anIntegerorRealby zero now raisesZeroDivisionError, as for Python numbers, instead ofValueError. - Behavior change: {class}
pikepdf.Matrixnow raisesTypeError, notValueError, for apikepdf.Objectthat is not a matrix or anObjectListwith a non-numeric element, matching what a native argument of the wrong type raises, so anexcept TypeErroraroundMatrix(*operands)behaves the same in either conversion mode. Size errors (must have 6 elements) remainValueError. - Behavior change:
bool()on apikepdf.Integerorpikepdf.Realis now by value (bool(pikepdf.Integer(0))isFalse), instead of raisingNotImplementedError: code is unreachable. - Behavior change: {meth}
~pikepdf.Object.as_dictand {meth}~pikepdf.Object.as_listnow accept adefaultargument and raiseTypeErroron a type mismatch, instead of raisingpikepdf.PdfErrorwith no way to supply a default. - Behavior change: Objects created in Python are now adopted by the
Pdfthey are inserted into, sopdf.Root.X = 42followed bypdf.Root.get_raw('/X').is_owned_by(pdf)isTrue, matching objects parsed from a file. Scalars (Name,String,Integer, ...) are adopted as a copy, so a handle held in a constant can still be inserted into several documents. ADictionaryorArrayis adopted in place and keeps its aliasing with the document; inserting the same container into a secondPdfnow raisesForeignObjectError, as it already did for containers parsed from a file. Construct a new object, or use {meth}pikepdf.Pdf.copy_foreign. Adoption is not permanent: removing the object from the document -- deleting or replacing the dictionary key or array element that held it, or destroying the document -- releases the claim, so the object can then be inserted into anotherPdf. If the same object was reachable under two keys, removing either one releases it; its value is unaffected and still readable through the remaining reference. - Behavior change: Objects copied into a document are adopted by it too.
The direct children of an object produced by
{meth}
pikepdf.Pdf.copy_foreign, {meth}pikepdf.Object.with_same_owner_as, {meth}pikepdf.Pdf.make_indirect, or by appending/inserting a page intoPdf.pages, and values inserted throughNameTree/NumberTree__setitem__, now report the destination document as their owner and follow itsconversion_mode. - Behavior change: {meth}
pikepdf.Object.with_same_owner_asand {meth}pikepdf.Pdf.make_indirectnow raiseForeignObjectErrorfor a direct object that already belongs to a differentPdf, instead of silently retagging an object that the other document still references. Remove it from that document first, or make it indirect there and use {meth}pikepdf.Pdf.copy_foreign. - Behavior change: {meth}
~pikepdf.Object.as_intwithcoerce=Truenow returns a supplied default for a value that does not fit in a 64-bit PDF integer, instead of raisingOverflowError. Called without a default it still raisesOverflowError. - Behavior change:
repr()of apikepdf.Objectnow honors the effective conversion mode of the object being displayed (context manager, then its owningPdf, then the global setting), rather than only the global setting. - Behavior change: {func}
pikepdf.set_object_conversion_modenow raisesValueErrorfor a value other than'implicit'or'explicit', instead of silently accepting it.
Performance
- {meth}
pikepdf.Pdf.savewrites directly to the file descriptor when saving to a filename or to a plain binary file object from {func}open, instead of calling the stream'swrite()method for every chunk of output. Saving a large document is about twice as fast. Other streams, such asBytesIO, pipes and subclasses of theioclasses, are written throughwrite()as before. isinstance()checks against {class}pikepdf.Dictionary, {class}pikepdf.Stream, {class}pikepdf.Integerand the other object classes are about five times faster.- {func}
pikepdf.unboxis now implemented in C++ and is about 40 times faster. {func}pikepdf.as_int, {func}pikepdf.as_floatand {func}pikepdf.as_decimalanswer for a nativeintorDecimalwithout constructing a PDF object, and converting a Real to aDecimalno longer imports thedecimalmodule each time.
Fixes
- {func}
pikepdf.models.metadata.decode_pdf_datenow accepts every PDF date form the specification allows, in which all fields after the year are optional.D:202001011230(no seconds) was misread as 12:03, andD:2020010112(no minutes) and offsets in hours only, such as-08', raisedValueError, so {attr}pikepdf.Pdf.docinfodates in these forms were not synchronized to XMP. - {meth}
pikepdf.Pdf.savenow raisesValueErrorwhen the destination is the stream thePdfwas opened from, or another stream on the same file, as it already did for the input file's path. Previously the new file was written into the input at the stream's current position, which produced a corrupt file and could corrupt the openPdf, since qpdf reads its input lazily. Open withallow_overwriting_input=Trueto save over the input. - Setting an XMP property that a file stored as plain text, such as
<dc:title>Old</dc:title>, now replaces the text. Previously the new value was added alongside the old text, which was still returned when the property was read back. - {attr}
pikepdf.PdfImage.filter_decodeparmsand {attr}pikepdf.PdfImage.decode_parmsno longer raiseAttributeErrorwhen an image's/DecodeParmsis a bare number or boolean, in either conversion mode. A/DecodeParmsthat is neither a dictionary nor an array is now ignored, as qpdf does, so each filter gets empty parameters. - An image
/Decodethat is not an array (a bare number, boolean, name or dictionary) is now ignored in favour of the default decode array, instead of raisingAttributeErrorwhen the image is read or converted. - {meth}
pikepdf.Stream.writenow raisesTypeErrorrather thanAttributeErrorwhenfilterordecode_parmsis a native Python value such as anint,strordict. - {meth}
pikepdf.Pdf.saveto an existing file that is not a regular file, such as/dev/null, a FIFO or a character device, now writes into it directly. Previously the temporary file was renamed over it, so when run as root, saving to/dev/nullreplaced the device with a regular file. - {meth}
pikepdf.Matrix.inversenow returns the correct translation for matrices whose determinant is not 1. Previously theeandfterms were not divided by the determinant, so the result was not a true inverse for any matrix that both scaled and translated. Matrices built only from rotations and translations were unaffected. (#743) - Storing a Python
intoutside the signed 64-bit range of a PDF integer, by assignment or inArray(...),Dictionary(...)orInteger(...), now raisesOverflowErrorinstead ofRuntimeError: std::bad_cast(or a misleadingTypeError). PdfImage,PdfImageBase,PdfJpxImage,PdfInlineImageandPaletteDataonce again report their module aspikepdf.models.image, the path they are imported from, instead of the privatepikepdf.models.image._classesintroduced when that module became a package in v10.10.- The examples for {func}
pikepdf.set_object_conversion_modeand {func}pikepdf.explicit_conversionno longer open a nonexistenttest.pdf, and the global-mode example restores the default mode when it is done. str()of apikepdf.Integer,BooleanorRealnow gives the value --'42','True','1.50'-- rather than the object's repr. In implicit mode a scalar arrives asint/bool/Decimalandstr()never reached the object, so every f-string, log line and string concatenation in code reading PDF values silently changed meaning under explicit mode.hash()now works onpikepdf.Integer,BooleanandReal(and on a null object), instead of raisingRuntimeError: don't know how to hash this. A scalar hashes like the Python value it compares equal to, soReal('1.0'),Decimal('1.0')and1agree, and a scalar can be a dict key or a set member as it could in implicit mode.- {meth}
pikepdf.Pdf.make_indirectand {meth}pikepdf.Object.with_same_owner_asno longer turn a direct scalar (Name,String,Integer,Real,Boolean,Operator) that you pass them into an indirect object in place. They make a copy indirect and return it, leaving your object direct and unowned. Previously the object you held silently became an indirect object of thatPdf, after whichhash()raised and any dict or set already holding it broke; a module-level constant such asFOO = Name.Foobecame tied to one document. Arrays, dictionaries and streams are still made indirect in place, so the object you passed stays aliased with the one in the document. - In explicit mode,
repr()of anArray,DictionaryorStreamnow names the scalars nested inside it --pikepdf.Real('42.42')rather than'42.42', which was indistinguishable from a PDF string and did not surviveeval(repr(obj)). - pikepdf's own higher-level APIs now work under explicit conversion mode.
They read values out of PDF objects and compute with them, and every such
read assumed the implicit-mode native type. Fixed in image extraction
(
PdfImageraisedNotImplementedErrorfor any indexed or /DeviceN colorspace), page labels (page.labelrestarted numbering at 1 and warned about a valid/St), outlines (a closed item read as open,flagscame back empty, and saving wrote the wrong/Count), form appearance generation (TypeErrorlaying out multiline and combed text fields),get_objects_with_ctm(a malformedcmoperator raised instead of being skipped),SimpleFontmetrics, andAction.new_window. PdfImagemetadata now reads the same in either conversion mode. A real/Widthor/Heightor a boolean/BitsPerComponentraisedNotImplementedErrorin explicit mode, and a real inside a/ColorSpacearray raised it in implicit mode.SimpleFont.ascent,descentandunscaled_char_width()now return aDecimalas documented, in either conversion mode. Font metrics stored as PDF integers were previously returned asint.- Iterating a
NamePath(e.g.list(path)) no longer falls back to the legacy__getitem__(0), (1), ...protocol, which never raisedIndexErrorand so iterated forever, exhausting memory. NamePath in objno longer silently answersFalsefor every path; see above.- Fixed a stale
NamePath['/A']['/B'].C[0]example in the type stub, which did not type-check against the runtime behavior (an instance's__getitem__only accepts anint). The correct form, matching the compiled docstring and the {doc}/topics/namepathdocumentation, isNamePath['/A']('/B').C[0]. - A direct object parsed from a file, then removed from the document (for
example
box = page.obj.get_raw('/MediaBox')followed bydel page.obj['/MediaBox']) and used after the document was closed, kept a dangling reference to the freed document and could crash the interpreter onrepr()or when inserted elsewhere. Removing an object from a document now releases its association with it, and aPdfdisconnects every object it adopted as it is destroyed.
Internals
- Linux wheel jobs now cache the compiled qpdf per build image with
actions/cache, so jobs that share an image no longer each download and compile the same qpdf. See the build process notes. - Fixed the Linux wheel build script's AlmaLinux detection, which was a malformed shell test that always evaluated false.
- The type stubs for the C++ extension module, until now a single 4,300-line
src/pikepdf/_core.pyi, are a stub-only packagesrc/pikepdf/_core/split into one stub per translation unit:_core/_matrix.pyicoverssrc/core/matrix.cpp,_core/_page.pyicoverssrc/core/page.cpp, and so on, with_core/__init__.pyire-exporting the lot and documenting the layout. Nothing changes at runtime or in the public API --pikepdf._coreis still one extension module andimport pikepdf._corestill resolves to it. Type checkers now name a type by its defining stub in messages (pikepdf._core._object.Objectrather thanpikepdf._core.Object);pikepdf.Objectremains the name to write in annotations. - API documentation for the C++ extension now lives only in the
src/pikepdf/_core/stubs, which are what Sphinx, type checkers and IDEs read. Docstrings attached to the C++ bindings were never published and had drifted from the stubs; they are merged into the stubs and removed from C++, and a test now fails if a binding gains a docstring or a public C++ name is missing from the stubs.help()on a C++ method now shows only its signature, as it already did for most of them.
Documentation
-
Restored the full {class}
pikepdf.Streamand {class}pikepdf.Dictionarydocumentation, including constructor arguments and examples, which had been reduced to a single line when those classes moved to C++. -
Clarified guidance on
Pdf.save(..., deterministic_id=)andstatic_id=, and the matchingJobBuilder.deterministic_id()andJobBuilder.static_id().deterministic_idgives reproducible, production-safe/IDvalues;static_idsets the same dummy/IDin every file and is for testing only.static_idnow appears last in thePdf.save()signature. Allsave()options are keyword-only, so existing code is unaffected.