Download Latest Version cccl-src-v3.4.3.zip (18.5 MB) Google Add to Preferred Sources
Home / python-1.2.1
Name Modified Size InfoDownloads / Week
Parent folder
CCCL Python Libraries (1.2.1) source code.tar.gz 2026-09-28 13.1 MB
CCCL Python Libraries (1.2.1) source code.zip 2026-09-28 21.6 MB
README.md 2026-09-28 2.6 kB
Totals: 3 Items   34.7 MB 0

CCCL Python Libraries (v1.2.1)

Previous release: v1.2.0.

These are the release notes for the cuda-cccl Python package version 1.2.1. This is a maintenance release with no new public Python APIs. The main changes are a self-contained Windows wheel, a workaround for a numba-cuda-mlir regression, clearer diagnostics for unsupported operations on structured data, and reduced Python-side overhead in cuda.compute calls.

Installation

Please refer to the install instructions here.

Packaging and dependencies

  • Windows wheels now bundle msvcp140.dll (#11495)

The Windows cuda-cccl wheel is repaired with delvewheel to include the MSVC C++ runtime DLL needed by cccl.c.parallel.dll. It no longer relies on whichever msvcp140.dll happens to be installed on the user's machine. The Windows cuda.compute test lanes now run without installing that DLL separately. Other packages, such as CuPy, may still have their own runtime requirements.

  • numba-cuda-mlir 0.5.3 excluded (#11617)

The cu12, cu13, sysctk12, and sysctk13 extras now require numba-cuda-mlir>=0.5.2,!=0.5.3,<0.6. Version 0.5.3 can fail when a stateful operator with the same state dtype and shape is compiled a second time in one process.

Bug fixes

  • Clearer errors for unsupported built-in operations on structs (#11404)

On the current V1 backend, using a built-in OpKind with a struct or other opaque storage type now raises a descriptive TypeError directing users to provide a custom operator, instead of failing later with a generic NVRTC compilation error. This does not change operations that use a built-in comparator on ordinary keys while carrying struct values.

Performance

  • Less per-call overhead in cuda.compute (#11497)

Shared iterator handling now caches whether an iterator represents a pointer and reuses its Pointer wrapper. Device-array pointer lookup also caches the appropriate accessor for each array type. These changes apply across algorithms without changing their public APIs or results. The PR measured approximately 9–20% less host-call overhead in representative reduce, scan, segmented-reduce, and histogram benchmarks; those measurements exclude GPU kernel execution.

Notes

  • v1.2.1 includes the Python-package changes since v1.2.0; there were no intervening cuda-cccl releases.
Source: README.md, updated 2026-09-28