| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-08-21 | 5.4 kB | |
| v2.9.0 source code.tar.gz | 2026-08-21 | 260.5 kB | |
| v2.9.0 source code.zip | 2026-08-21 | 311.7 kB | |
| Totals: 3 Items | 577.6 kB | 0 | |
oneAPI v2.9.0
Highlights: work is now submitted through a per-task immediate command list instead of a throwaway command list per launch, copy and fill. This removes the per-dispatch driver garbage that pushed NEO into allocation failure under launch storms, and cuts submission overhead. Alongside it, a scratch hedge makes the driver's one unrecoverable allocation — the scratch buffer it allocates on the first submission of a spilling kernel — happen at a clean moment instead of at a GC-lottery-determined one. The minimum driver requirement tightens to one implementing in-order immediate command lists, and
global_queueno longer refers to the submission path; see the breaking changes.
Highlights
- Per-task
oneStream: every kernel launch,copyto!andfill!is appended to one in-order asynchronous immediate command list, submitted to the device as it happens. Previously each dispatch created, executed and dropped its own command list, leaving the destruction of the driver objects behind it (lists, command buffers, heaps) to finalizer timing; that object class no longer exists. The stream is reachable through the new exportedglobal_stream(ctx, dev);synchronize()andsynchronize(::oneStream)drain it. (#610) - Explicit ordering at the oneMKL boundary. oneMKL still needs a real command queue for SYCL interop, so each stream lazily creates a companion queue — a separate execution stream from the driver's point of view.
sycl_queuedrains the immediate list before handing the queue out (Julia → MKL), and a dirty flag makes the next Julia-side submission wait on the queue (MKL → Julia; oneBoolload on the fast path). FFT plans, which capture their queue at construction, apply the boundary in their_exec!methods. Interleave tests (broadcast → gemm/fft → broadcast with no intermediate synchronization) cover both directions. (#610) - Scratch hedge. NEO allocates a stream's scratch buffer at the first submission of a kernel whose register spill exceeds what is already allocated, and that allocation has no error path: when it fails the process aborts inside the submit call (
UNRECOVERABLE_IF, identical from the 25.18 LTS driver through current master), with no Julia-side recourse.ZeKernelnow caches its spill size and each stream tracks a spill high-water mark; before the first submission that crosses it, the launch path retires in-flight work, flushes deferred releases and runsGC.gc(false), so the allocation happens with the least possible dead-but-unfinalized driver memory live. Fires once per (stream, spill tier) — one synchronize per workload in practice — and costs anIntcompare otherwise. Opt out withONEAPI_SCRATCH_HEDGE=0. (#609, [#610]) - Support for immediate lists is probed by creating one, not by the reported API version: the Aurora LTS driver reports Level Zero 1.6 while fully implementing in-order immediate lists (they are DPC++'s production submission path on PVC). For the same reason no 1.9-only loader entrypoints are used. A driver that rejects the creation gets a clear error at first launch instead of a fallback. (#610)
- LTS drain-before-free follows the new shape: the command-queue registry becomes a stream registry, and both the immediate list and the companion queue of every registered stream are drained before a buffer with possibly in-flight work is freed. Immediate lists get the same bounded-timeout finalizer drain as queues.
KernelAbstractions.priority!now swaps the task's stream for one created with the requested priority. (#610) ONEAPI_SYNC_EACH_SUBMISSION=1now host-synchronizes the stream after every append. The dropped-tail driver bug it works around was only ever observed on the queue-submission path; the knob is kept until its absence on immediate lists is re-established under oversubscription. (#610)- CI: failed tests are retried (#617); fork PRs get the Runic check and the ALCF Aurora pipeline again (
actions/checkout@v7fork opt-in; ALCF pipelines created through a trigger token so Jacamar no longer refuses the bot attribution). (#617, [#620])
Breaking changes
- The driver must support in-order immediate command lists (Level Zero ≥ 1.9, or a driver implementing them regardless of its reported version, as Aurora's LTS stack does). There is no fallback submission path: on a driver that rejects the creation, the first launch errors with
oneAPI.jl requires driver support for in-order immediate command lists …. (#610) global_queue(ctx, dev)now returns the stream's lazily-created oneMKL companion queue, which is not where@oneapi,copyto!andfill!submit. In particularsynchronize(global_queue(ctx, dev))no longer waits for launched kernels or copies; usesynchronize()orsynchronize(global_stream(ctx, dev)).@oneapi queue=...with an explicit queue still works and still submits through a per-dispatch command list. (#610)- The internal
synchronize_all_queuesis renamedsynchronize_all_streams. (#610)
Merged pull requests:
- Guard NEO's scratch allocation with a spill-triggered drain (#609) (@michel2323)
- Submit work through per-task immediate command lists (#610) (@michel2323)
- Retry test failures (#617) (@christiangnrd)
- Fix fork-PR CI: Runic checkout and ALCF pipeline attribution (#620) (@michel2323)