Download Latest Version v2.9.2 source code.zip (315.7 kB) Google Add to Preferred Sources
Home / v2.9.0
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-08-21 5.4 kB
v2.9.0 source code.tar.gz 2026-08-21 260.5 kB
v2.9.0 source code.zip 2026-08-21 311.7 kB
Totals: 3 Items   577.6 kB 0

oneAPI v2.9.0

Diff since v2.8.2

Highlights: work is now submitted through a per-task immediate command list instead of a throwaway command list per launch, copy and fill. This removes the per-dispatch driver garbage that pushed NEO into allocation failure under launch storms, and cuts submission overhead. Alongside it, a scratch hedge makes the driver's one unrecoverable allocation — the scratch buffer it allocates on the first submission of a spilling kernel — happen at a clean moment instead of at a GC-lottery-determined one. The minimum driver requirement tightens to one implementing in-order immediate command lists, and global_queue no longer refers to the submission path; see the breaking changes.

Highlights

  • Per-task oneStream: every kernel launch, copyto! and fill! is appended to one in-order asynchronous immediate command list, submitted to the device as it happens. Previously each dispatch created, executed and dropped its own command list, leaving the destruction of the driver objects behind it (lists, command buffers, heaps) to finalizer timing; that object class no longer exists. The stream is reachable through the new exported global_stream(ctx, dev); synchronize() and synchronize(::oneStream) drain it. (#610)
  • Explicit ordering at the oneMKL boundary. oneMKL still needs a real command queue for SYCL interop, so each stream lazily creates a companion queue — a separate execution stream from the driver's point of view. sycl_queue drains the immediate list before handing the queue out (Julia → MKL), and a dirty flag makes the next Julia-side submission wait on the queue (MKL → Julia; one Bool load on the fast path). FFT plans, which capture their queue at construction, apply the boundary in their _exec! methods. Interleave tests (broadcast → gemm/fft → broadcast with no intermediate synchronization) cover both directions. (#610)
  • Scratch hedge. NEO allocates a stream's scratch buffer at the first submission of a kernel whose register spill exceeds what is already allocated, and that allocation has no error path: when it fails the process aborts inside the submit call (UNRECOVERABLE_IF, identical from the 25.18 LTS driver through current master), with no Julia-side recourse. ZeKernel now caches its spill size and each stream tracks a spill high-water mark; before the first submission that crosses it, the launch path retires in-flight work, flushes deferred releases and runs GC.gc(false), so the allocation happens with the least possible dead-but-unfinalized driver memory live. Fires once per (stream, spill tier) — one synchronize per workload in practice — and costs an Int compare otherwise. Opt out with ONEAPI_SCRATCH_HEDGE=0. (#609, [#610])
  • Support for immediate lists is probed by creating one, not by the reported API version: the Aurora LTS driver reports Level Zero 1.6 while fully implementing in-order immediate lists (they are DPC++'s production submission path on PVC). For the same reason no 1.9-only loader entrypoints are used. A driver that rejects the creation gets a clear error at first launch instead of a fallback. (#610)
  • LTS drain-before-free follows the new shape: the command-queue registry becomes a stream registry, and both the immediate list and the companion queue of every registered stream are drained before a buffer with possibly in-flight work is freed. Immediate lists get the same bounded-timeout finalizer drain as queues. KernelAbstractions.priority! now swaps the task's stream for one created with the requested priority. (#610)
  • ONEAPI_SYNC_EACH_SUBMISSION=1 now host-synchronizes the stream after every append. The dropped-tail driver bug it works around was only ever observed on the queue-submission path; the knob is kept until its absence on immediate lists is re-established under oversubscription. (#610)
  • CI: failed tests are retried (#617); fork PRs get the Runic check and the ALCF Aurora pipeline again (actions/checkout@v7 fork opt-in; ALCF pipelines created through a trigger token so Jacamar no longer refuses the bot attribution). (#617, [#620])

Breaking changes

  • The driver must support in-order immediate command lists (Level Zero ≥ 1.9, or a driver implementing them regardless of its reported version, as Aurora's LTS stack does). There is no fallback submission path: on a driver that rejects the creation, the first launch errors with oneAPI.jl requires driver support for in-order immediate command lists …. (#610)
  • global_queue(ctx, dev) now returns the stream's lazily-created oneMKL companion queue, which is not where @oneapi, copyto! and fill! submit. In particular synchronize(global_queue(ctx, dev)) no longer waits for launched kernels or copies; use synchronize() or synchronize(global_stream(ctx, dev)). @oneapi queue=... with an explicit queue still works and still submits through a per-dispatch command list. (#610)
  • The internal synchronize_all_queues is renamed synchronize_all_streams. (#610)

Merged pull requests:

  • Guard NEO's scratch allocation with a spill-triggered drain (#609) (@michel2323)
  • Submit work through per-task immediate command lists (#610) (@michel2323)
  • Retry test failures (#617) (@christiangnrd)
  • Fix fork-PR CI: Runic checkout and ALCF pipeline attribution (#620) (@michel2323)
Source: README.md, updated 2026-08-21