Download Latest Version v0.1.15 source code.zip (15.3 MB) Google Add to Preferred Sources
Home / v0.1.15
Name Modified Size InfoDownloads / Week
Parent folder
tilelang-0.1.15-cp39-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl 2026-09-30 49.6 MB
tilelang-0.1.15.tar.gz 2026-09-30 93.3 MB
tilelang-0.1.15-cp39-abi3-macosx_11_0_arm64.whl 2026-09-30 33.7 MB
tilelang-0.1.15-cp39-abi3-win_amd64.whl 2026-09-30 29.5 MB
tilelang-0.1.15-cp39-abi3-manylinux_2_34_aarch64.whl 2026-09-30 44.8 MB
README.md 2026-09-29 17.3 kB
v0.1.15 source code.tar.gz 2026-09-29 14.1 MB
v0.1.15 source code.zip 2026-09-29 15.3 MB
Totals: 8 Items   280.3 MB 0

Highlights

  • Native Huawei Ascend 950 support (#3308): an end-to-end NPU backend with native code generation, automatic Cube/Vector scheduling and synchronization, and mixed SIMD/SIMT programming.
  • Automatic CUDA warp specialization (#3059, [#3185]): an opt-in role-based scheduler that assigns TMA loads, MMA computation, TMA stores, and worker operations to specialized warp groups.
  • Unified block-scaled GEMM (#3237, [#3257], [#3284]): common T.gemm_blockscaled semantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support.
  • More expressive Python frontend (#3230): compile-time iteration over Python iterables, enumerate, zip, comprehensions, and generator expressions.

Ascend 950

  • Add tilelang.ascend.language and target="ascend" for Huawei Ascend 950 (dav-3510).
  • Combine Cube GEMM and Vector computation in one kernel, with T.SimdVF and T.SimtVF regions.
  • Support explicit UB/L1/L0 storage, tiled copies, cross-core transfers, and MXFP8/MXFP4 block-scaled GEMM.
  • Add automatic scheduling, pipelining, multi-buffering, layout inference, and synchronization insertion.
  • Integrate Bisheng compilation, tvm_ffi and Cython execution, PyTorch NPU tensors and streams, and NPU profiling.
  • Include GEMM, DeepGEMM-style kernels, FlashAttention forward/backward, RMSNorm, and FP8 quantization examples.

See the Ascend 950 guide for installation and usage. Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects.

CUDA

  • Enable role-based automatic warp specialization with TL_ENABLE_AUTO_WARP_SPECIALIZATION: "role_based"; add GEMM and FlashAttention examples (#3059, [#3185]).
  • Extend SM120 block-scaled GEMM to fragment-resident A operands, row-major scale fragments, and odd per-warp atom grids; diagnose conflicting scale-fragment layouts (#3257, [#3284]).
  • Add round-to-nearest T.fma and T.fmul, and complete additional FP16/BF16 math bridges (#3134, [#3132], [#3163]).
  • Preserve packed FP8 vector copies and fuse exact FP4-to-FP8 conversions through FP32 (#3276, [#3204]).
  • Fix WGMMA/UMMA K-panel strides for operand layouts and insert async-proxy fences for sparse MMA (#2965, [#3139]).
  • Correct atomic vectorization for invariant or non-contiguous destinations; keep shared FP32 atomics scalar on SM90 (#3129, [#3219], [#3238]).
  • Fix NVRTC warp reductions, kernel-body assertions, boolean bitwise negation, and non-constant int4/uint4 broadcasts; restrict 256-bit global loads/stores to SM100+ (#3260, [#3206], [#3228], [#3114], [#3248]).
  • Add an FP8 sparse MLA forward example for DeepSeek V3.2 on Hopper (#3224).

Language and Compiler

  • Make kernel launch encoding target-neutral, with launch options and operation hints owned by their backend dialects (#3186, [#3203]).
  • Support Python compile-time iteration and comprehensions while preserving device-loop semantics for for ... in range(...) (#3230).
  • Support unary plus on symbolic expressions, fix loop-variable binding, remove warnings for immutable rebinding, and reject unsupported loop else clauses (#3141, [#3232], [#3262], [#3142]).
  • Treat T.copy / T.async_copy coalesced_width as a hint and clamp it to the achievable vector width (#3246).
  • Preserve dynamic reduction tail guards, fix parallel-loop lowering with let inlining disabled, and restore async-copy lowering with partitioned layouts (#3294, [#3269], [#3278]).
  • Improve reducer layout planning and symbolic layout validation; reject unsupported AllReduce thread strides and reduction NaN-propagation dtypes (#3171, [#3233], [#3266], [#3273]).
  • Fix NaN-propagating clamp lowering and eliminate unused bindings during simplification (#3205, [#3293]).

ROCm, CPU, and Metal

  • ROCm: support the DeepSeek V3.2 Top-K selector, wave64 Hadamard transforms, and GLM-5.3 k-pool examples, including Top-K index transformation (#3147, [#3154], [#3254]).
  • Expand ROCm validation for attention kernels and portable examples, and document pip installation (#3146, [#3148], [#3165], [#3168]).
  • CPU: compute scalar GEMM products in the accumulator dtype and legalize BF16 arithmetic (#2917, [#3201]).
  • Metal: support 32-bit integer atomic add and respect GEMM buffer-region offsets (#3211, [#3209]).

Runtime, Build, and Tooling

  • Store cached kernel parameters as JSON instead of cloudpickle (#3143).
  • Publish CUDA binaries and metadata atomically in immutable cache directories, preventing readers from observing partially published entries (#3177).
  • Fix dynamic-output allocation when the sizing input follows the output, and add missing uint64 argument mappings (#3207, [#3229]).
  • Add wall-clock benchmarking and MPS-compatible timing helpers (#3234).
  • Improve CUDA/ROCm compiler discovery and library loading for symlinked installations (#2839, [#3166], [#3227]).
  • Enable optimization for default single-config native builds and restore effective Windows wheel-build caching (#3191, [#3133], [#3305]).
  • Validate built wheels on GPU runners and make performance regressions fail CI (#3167, [#3182]).

Compatibility Notes

  • Remove the unused sync and group parameters from T.Pipelined, and k_pack from T.gemm_sp (#3202).
  • ROCm kernels using T.gemm(k_pack=...) should import tilelang.rocm.language (#3203).
  • Kernels querying thread extents during tracing must specify threads= explicitly in T.Kernel (#3186).
  • T.symbolic remains available as a deprecated alias; use T.dynamic for new code (#3216).
  • Remove TILELANG_CACHE_VERIFY_HASH; binary artifact hash verification is now mandatory. Legacy cache formats are rebuilt automatically (#3143, [#3177]).

Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.14...v0.1.15

What's Changed

New Contributors

Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.14...v0.1.15

Source: README.md, updated 2026-09-29