Download Latest Version v0.1.15 source code.zip (15.3 MB) Google Add to Preferred Sources
Home / v0.1.14
Name Modified Size InfoDownloads / Week
Parent folder
tilelang-0.1.14.tar.gz 2026-09-02 91.7 MB
tilelang-0.1.14-cp39-abi3-win_amd64.whl 2026-09-02 28.5 MB
tilelang-0.1.14-cp39-abi3-manylinux_2_34_aarch64.whl 2026-09-02 41.0 MB
tilelang-0.1.14-cp39-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-02 45.4 MB
tilelang-0.1.14-cp39-abi3-macosx_11_0_arm64.whl 2026-09-02 32.7 MB
README.md 2026-09-01 24.8 kB
v0.1.14 source code.tar.gz 2026-09-01 12.6 MB
v0.1.14 source code.zip 2026-09-01 13.6 MB
Totals: 8 Items   265.5 MB 0

Highlights

  • Reducer v2 (#2940, [#3093], [#3043], [#3044], [#3079], [#3100]): T.alloc_reducer reworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates.
  • Warp specialization schedules (#2892): new scheduling and materialization mechanism for warp-specialized kernels.
  • Layout inference cost models (#2960, [#3055], [#3061]): new IO-aware cost model for free-mode layout selection; register-count restored as the default, with an environment override to switch models.
  • Unified backend resolution policy (#2318) plus backend split-up (#2855, [#2870], [#2850]): backend selection is now resolved through a single policy, and builtin ops / Python op proxies are split per backend (CUDA/ROCm/Metal).
  • Compilation speed: up to ~4x faster cold parallel/AOT compilation (#2809); Z3 solvers materialized lazily (#3105) and analyzer contexts isolated per kernel compilation (#2890).
  • TMA rework: TMA copy lowering unified on CuTe algebra (#3106); TMA layouts made region-aware to keep slices contiguous (#3089).

Language

  • Recycle T.unroll(explicit=True) for early explicit unrolling (#2859)
  • Make the region bridge a builtin intrinsic (#2983)
  • Expose cluster_mask on T.tma_copy (#2932)
  • Unify contiguous stride construction under a single implementation (#3016); honor declared strides in pointer helpers (#3073)
  • Stricter validation: reject symbolic T.gemm tile dimensions with a clear message (#3113), validate T.gemm k_pack arguments (#3094), reject non-positive arrive_count in alloc_barrier/alloc_cluster_barrier (#3112), reject T.Parallel indexing of local buffers (#3041), reject break in fully expanded loops (#3078)

CUDA

  • tcgen05: pack logical TMEM buffers into shared tcgen05.alloc arenas (#2831); support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st (#2880); fix ld/st segment pointer advancement in b32 columns (#2952)
  • Select the widest legal WGMMA N instead of gcd (#2931)
  • FP32x2 accumulation for reductions: per-reduce control (#3057) and a global PassConfig (#3128)
  • Pre-SM80 fallback for bf16 atomic add (#2938); int4x2/uint4x2 codegen (#3036); 16-bit CUTLASS type overloads for fast-math, __ldg, and htan intrinsics (#3097, [#3077], [#3028], [#2894])
  • Fixes: warp shuffle for half/bfloat16/FP8 (#3056), FP8 min/max codegen (#3047), UB in packed 8-bit vector stores (#3092), logical not for vectorized bool (#3117, also HIP), vectorized Select codegen (#2843), ldmatrix source offsets wrapped within shared-memory regions (#3110), NVRTC kernel handles isolated per adapter (#2950), flat CUDA include discovery for NVRTC (#2829), masked warpsync in in-warp allreduce (#2865)
  • Remove T.{reads,writes} for T.tma_{gather4,scatter4} (#3053); separate TMA atomic-add dtype support from layout encoding (#2846)

ROCm and other backends

  • Remove the Composable Kernel dependency (#3111); ROCm CI re-enabled on a gfx942 runner (#2874, [#2910])
  • Fixes: preserve FP8 bits in warp shuffles (#3104), lower vector Select conditions lane-wise (#2889), emit a compiler barrier for tl.sync_warp on HIP (#2872), reject sub-wavefront block sizes instead of crashing (#2918), resolve versioned device properties in the HIP stub (#2919); emit #line directives for the HIP target (#3058)
  • CPU backend: support atomic ops (#2941) and reduce ops (#2893)
  • Metal: preserve pointer address spaces for byte offsets (#2925); resolve auto backend to torch and skip disk cache for torch (#2856)
  • CuTeDSL: port backend intrinsics to CUTLASS DSL primitives (#2871)

Compiler / Transform

  • Refactor the loop vectorization plan with ConstraintKind (#2935); always vectorize T.Parallel loops (#3121); scalarize Select in automatic vectorization (#3060)
  • Add VerifyBufferInit, a general buffer-initialization check (#2956)
  • Debug info: preserve source spans across lowering passes (#2966); emit #line directives from TIR spans (#3048)
  • Fixes: don't drop syncs from the other if branch (#3085), fix wait parity for explicit mbarriers in pipelined loops (#3087), avoid int32 overflow in vector analysis (#3066), fix non-divisible nested modulo simplification (#3065), fix ties-away-from-zero round compile on bfloat16/float8 (#2873), fix absmax/abssum for uint dtypes (#2845), carry memory_order through vectorized atomic_add (#2924), fix unsigned zero-point decode underflow (#3118), keep cp.async operands in their address spaces (#2869), handle grid barriers and unbounded pointer ranges (#3050), deduplicate replicated reducer updates (#2881), bind symbolic coordinate ranges in FragmentThreadIndexProbe (#3096), reject non-round-tripping inferred layout inverses (#3090)
  • Cleanup: remove obsolete compiler and runtime paths (#3086); remove the obsolete disable-fast-math pass config (#3098)

Runtime / JIT / Build

  • Kernel cache: detect and repair corrupted cache entries (#3074), export libraries after disk-cache hits (#3116), remove the separate cache temporary directory (#3069)
  • Allocate kernel outputs through the packed API (#2937); fix get_parent_locals frame self-reference leak (#2934)
  • Support Cython 3.3 with the Python 3.9 limited API (#3068); fix CMake reconfigure aborting in FindPipCUDAToolkit before project() (#3102)

Tooling / Ecosystem

  • Official compile-only CLI (#3045)
  • Unified pass instrumentation per compilation (#2923); Pass Visualizer driven by PassInstrument (#2866); show kernel name in the pass timing report (#2905)
  • Data race check disabled by default, opt-in via env var (#2851)
  • Open-source TileLang LSP announced (#2862); new agent skills: simplification (#3080), semantic validation (#3054), PR submission (#3082), backend architecture (#2900)
  • Examples: generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling (#2902); adopt multi-staged buffers in examples (#2836)
  • Docs: add Sunrise-AI TANG (#3107), HYGON (#3101), and MetaX MACA (#3095) to supported platforms; link the multi-backend architecture design (#3123)

What's Changed

New Contributors

Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.13...v0.1.14

Source: README.md, updated 2026-09-01