| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| tilelang-0.1.15-cp39-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl | 2026-09-30 | 49.6 MB | |
| tilelang-0.1.15.tar.gz | 2026-09-30 | 93.3 MB | |
| tilelang-0.1.15-cp39-abi3-macosx_11_0_arm64.whl | 2026-09-30 | 33.7 MB | |
| tilelang-0.1.15-cp39-abi3-win_amd64.whl | 2026-09-30 | 29.5 MB | |
| tilelang-0.1.15-cp39-abi3-manylinux_2_34_aarch64.whl | 2026-09-30 | 44.8 MB | |
| README.md | 2026-09-29 | 17.3 kB | |
| v0.1.15 source code.tar.gz | 2026-09-29 | 14.1 MB | |
| v0.1.15 source code.zip | 2026-09-29 | 15.3 MB | |
| Totals: 8 Items | 280.3 MB | 0 | |
Highlights
- Native Huawei Ascend 950 support (#3308): an end-to-end NPU backend with native code generation, automatic Cube/Vector scheduling and synchronization, and mixed SIMD/SIMT programming.
- Automatic CUDA warp specialization (#3059, [#3185]): an opt-in role-based scheduler that assigns TMA loads, MMA computation, TMA stores, and worker operations to specialized warp groups.
- Unified block-scaled GEMM (#3237, [#3257], [#3284]): common
T.gemm_blockscaledsemantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support. - More expressive Python frontend (#3230): compile-time iteration over Python iterables,
enumerate,zip, comprehensions, and generator expressions.
Ascend 950
- Add
tilelang.ascend.languageandtarget="ascend"for Huawei Ascend 950 (dav-3510). - Combine Cube GEMM and Vector computation in one kernel, with
T.SimdVFandT.SimtVFregions. - Support explicit UB/L1/L0 storage, tiled copies, cross-core transfers, and MXFP8/MXFP4 block-scaled GEMM.
- Add automatic scheduling, pipelining, multi-buffering, layout inference, and synchronization insertion.
- Integrate Bisheng compilation,
tvm_ffiand Cython execution, PyTorch NPU tensors and streams, and NPU profiling. - Include GEMM, DeepGEMM-style kernels, FlashAttention forward/backward, RMSNorm, and FP8 quantization examples.
See the Ascend 950 guide for installation and usage. Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects.
CUDA
- Enable role-based automatic warp specialization with
TL_ENABLE_AUTO_WARP_SPECIALIZATION: "role_based"; add GEMM and FlashAttention examples (#3059, [#3185]). - Extend SM120 block-scaled GEMM to fragment-resident A operands, row-major scale fragments, and odd per-warp atom grids; diagnose conflicting scale-fragment layouts (#3257, [#3284]).
- Add round-to-nearest
T.fmaandT.fmul, and complete additional FP16/BF16 math bridges (#3134, [#3132], [#3163]). - Preserve packed FP8 vector copies and fuse exact FP4-to-FP8 conversions through FP32 (#3276, [#3204]).
- Fix WGMMA/UMMA K-panel strides for operand layouts and insert async-proxy fences for sparse MMA (#2965, [#3139]).
- Correct atomic vectorization for invariant or non-contiguous destinations; keep shared FP32 atomics scalar on SM90 (#3129, [#3219], [#3238]).
- Fix NVRTC warp reductions, kernel-body assertions, boolean bitwise negation, and non-constant int4/uint4 broadcasts; restrict 256-bit global loads/stores to SM100+ (#3260, [#3206], [#3228], [#3114], [#3248]).
- Add an FP8 sparse MLA forward example for DeepSeek V3.2 on Hopper (#3224).
Language and Compiler
- Make kernel launch encoding target-neutral, with launch options and operation hints owned by their backend dialects (#3186, [#3203]).
- Support Python compile-time iteration and comprehensions while preserving device-loop semantics for
for ... in range(...)(#3230). - Support unary plus on symbolic expressions, fix loop-variable binding, remove warnings for immutable rebinding, and reject unsupported loop
elseclauses (#3141, [#3232], [#3262], [#3142]). - Treat
T.copy/T.async_copycoalesced_widthas a hint and clamp it to the achievable vector width (#3246). - Preserve dynamic reduction tail guards, fix parallel-loop lowering with let inlining disabled, and restore async-copy lowering with partitioned layouts (#3294, [#3269], [#3278]).
- Improve reducer layout planning and symbolic layout validation; reject unsupported AllReduce thread strides and reduction NaN-propagation dtypes (#3171, [#3233], [#3266], [#3273]).
- Fix NaN-propagating clamp lowering and eliminate unused bindings during simplification (#3205, [#3293]).
ROCm, CPU, and Metal
- ROCm: support the DeepSeek V3.2 Top-K selector, wave64 Hadamard transforms, and GLM-5.3 k-pool examples, including Top-K index transformation (#3147, [#3154], [#3254]).
- Expand ROCm validation for attention kernels and portable examples, and document pip installation (#3146, [#3148], [#3165], [#3168]).
- CPU: compute scalar GEMM products in the accumulator dtype and legalize BF16 arithmetic (#2917, [#3201]).
- Metal: support 32-bit integer atomic add and respect GEMM buffer-region offsets (#3211, [#3209]).
Runtime, Build, and Tooling
- Store cached kernel parameters as JSON instead of cloudpickle (#3143).
- Publish CUDA binaries and metadata atomically in immutable cache directories, preventing readers from observing partially published entries (#3177).
- Fix dynamic-output allocation when the sizing input follows the output, and add missing
uint64argument mappings (#3207, [#3229]). - Add wall-clock benchmarking and MPS-compatible timing helpers (#3234).
- Improve CUDA/ROCm compiler discovery and library loading for symlinked installations (#2839, [#3166], [#3227]).
- Enable optimization for default single-config native builds and restore effective Windows wheel-build caching (#3191, [#3133], [#3305]).
- Validate built wheels on GPU runners and make performance regressions fail CI (#3167, [#3182]).
Compatibility Notes
- Remove the unused
syncandgroupparameters fromT.Pipelined, andk_packfromT.gemm_sp(#3202). - ROCm kernels using
T.gemm(k_pack=...)should importtilelang.rocm.language(#3203). - Kernels querying thread extents during tracing must specify
threads=explicitly inT.Kernel(#3186). T.symbolicremains available as a deprecated alias; useT.dynamicfor new code (#3216).- Remove
TILELANG_CACHE_VERIFY_HASH; binary artifact hash verification is now mandatory. Legacy cache formats are rebuilt automatically (#3143, [#3177]).
Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.14...v0.1.15
What's Changed
- [CI] Cap torch<2.14 for CUDA tests until flash-attn supports torch 2.14 by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3136
- Bump transformers from 5.5.0 to 5.10.1 in /examples/bitnet-1.58b by @dependabot[bot] in https://github.com/tile-ai/tilelang/pull/3131
- [CI] Fix slow Windows wheel builds by making ccache effective by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3133
- [Language][CUDA] Add T.fma and T.fmul round-to-nearest intrinsics by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3134
- [CUDA] Complete 16-bit bridges for CUDA-lowered unary math by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3132
- [Refactor] Restore LoopUnswitching test altered by [#3121] by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3124
- [BugFix][CUDA] Support non-constant int4/uint4 broadcast by @jjppp in https://github.com/tile-ai/tilelang/pull/3114
- [BugFix][CPU] Compute scalar GEMM products in accum dtype instead of input dtype by @Dino1844 in https://github.com/tile-ai/tilelang/pull/2917
- [BugFix][Hopper][Blackwell] Read the WGMMA/UMMA K-panel stride from the layout instead of the operand extent by @bigSheep123 in https://github.com/tile-ai/tilelang/pull/2965
- [BugFix][Carver] Gate _legalize_info fallback on sm_version by @mocusez in https://github.com/tile-ai/tilelang/pull/3145
- [ROCm] Validate block-causal attention kernels by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3146
- [ROCm] Validate attention sink kernels by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3148
- [ROCm] Support DeepSeek-V3.2 Top-K selector by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3147
- [Bugfix][Language] Reject loop else clauses in the eager frontend by @rishabhsinha17 in https://github.com/tile-ai/tilelang/pull/3142
- [Example] Fix hadamard reference precision on TF32 and add pytest entry for CI by @Guan-jeans in https://github.com/tile-ai/tilelang/pull/3126
- [BugFix][Carver] Compare sm_version numerically in plan_rasterization by @mocusez in https://github.com/tile-ai/tilelang/pull/3144
- [Misc] Expose .agents skills to Claude Code via .claude/skills symlinks by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3161
- [Language] Cache dtype.as_torch and demote storage-dtype fallback logs to debug by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3160
- [Refactor][CUDA] Replace per-lane-count vector constructors with variadic packers by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3158
- [CUDA] Add 16-bit hpow and hfmod bridges by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3163
- [BugFix][ROCm] Resolve hipcc robustly instead of trusting bare PATH lookup by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3166
- [CI] Validate built wheels on GPU runners in the Dist workflow by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3167
- [Doc] Document pip installation on AMD GPUs (ROCm) by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3168
- [Bugfix][Language] Support unary plus on PrimExpr in the eager frontend by @rishabhsinha17 in https://github.com/tile-ai/tilelang/pull/3141
- [CUDA] Add async-proxy fences for sparse MMA intrinsics by @ZenAlexa in https://github.com/tile-ai/tilelang/pull/3139
- [Feature] [CUDA] Role-based automatic warp specialization by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3059
- [CUDA] Resolve symlinked nvcc before deriving CUDA_HOME by @morluto in https://github.com/tile-ai/tilelang/pull/2839
- [Language] Remove deprecated T.symbolic alias by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3122
- [ROCm] Support wave64 Hadamard transforms by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3154
- [Cache] Store cached kernel params as JSON instead of cloudpickle by @rishabhsinha17 in https://github.com/tile-ai/tilelang/pull/3143
- [BugFix][Transform] Keep shared FP32 atomics scalar on SM90 by @ZenAlexa in https://github.com/tile-ai/tilelang/pull/3129
- [Layout] Consider scalar reducer plans in register-count search by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3171
- [Transform] Add snapshot mode to FreshenMutableReads by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3174
- [CUDA][Cache] Publish immutable binary cache directories by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3177
- [CI] [pre-commit.ci] autoupdate by @pre-commit-ci[bot] in https://github.com/tile-ai/tilelang/pull/3181
- [CI] Fail the perf bot on a regression instead of only reporting it by @cklxx in https://github.com/tile-ai/tilelang/pull/3182
- [Refactor] [CUDA] Rename
AutoScheduletoAutoWarpSpecializationby @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3185 - [Refactor][Language] Make T.Kernel target-neutral with dialect-owned launch annotations by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3186
- [Cleanup][Language] Remove dead T.Pipelined(sync/group) and T.gemm_sp(k_pack) parameters by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3202
- [Refactor][Language] Move backend-specific op hints into their owning dialects by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3203
- [BugFix][CPU] Legalize BF16 arithmetic in the CPU pipeline by @anerli in https://github.com/tile-ai/tilelang/pull/3201
- [Language] Restore deprecated T.symbolic alias by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3216
- [Refactor][JIT] Replace callee-allocated-output target sniffing with a backend capability flag by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3218
- [Transform] Restore bind-before-use order when merging constraint sets by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3221
- [BugFix][Metal] Respect GEMM buffer region offsets by @anerli in https://github.com/tile-ai/tilelang/pull/3209
- [Runtime] Fix library loading for symlink installs by @sepcnt in https://github.com/tile-ai/tilelang/pull/3227
- [CUDA] Fix boolean bitwise negation codegen by @sepcnt in https://github.com/tile-ai/tilelang/pull/3228
- [Example] FP8 sparse MLA forward for DeepSeek V3.2 on Hopper by @xuebozhang525-alt in https://github.com/tile-ai/tilelang/pull/3224
- [Profiler] Split timing helpers and add wall-clock benchmarking by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3234
- [CPU] Import the CPU dialect in CPU tests by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3236
- [BugFix][Vectorize] Keep atomic_add scalar for invariant/non-contiguous destinations by @Dino1844 in https://github.com/tile-ai/tilelang/pull/3219
- [Transform][CUDA] Plan atomic vector widths from destination addresses by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3238
- [Do not review][Op][Language][CUDA] Add common block-scaled GEMM semantics and backend dispatch by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3237
- [JIT] Add missing uint64 argument type mappings by @sepcnt in https://github.com/tile-ai/tilelang/pull/3229
- [BugFix] Restore symbolic loop-layout injectivity proof; reject layouts on symbolic shared tiles by @sepcnt in https://github.com/tile-ai/tilelang/pull/3233
- [ROCm] Run portable example validation in CI by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3165
- [BugFix][CUDA] Only emit 256-bit global load/store on SM100+ targets by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3248
- [Metal] Support 32-bit integer atomic add by @anerli in https://github.com/tile-ai/tilelang/pull/3211
- [Testing] Drop duplicated codegen-smoke tests; fix fastmath self-comparison assertions by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3252
- [ROCm] Add GLM-5.3 k-pool Top-K transform by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3254
- [Frontend] Support Python iterables and comprehensions by @sepcnt in https://github.com/tile-ai/tilelang/pull/3230
- [Frontend] Remove warnings for immutable variable rebinding by @LJC00118 in https://github.com/tile-ai/tilelang/pull/3262
- [BugFix] Reject non-power-of-two AllReduce thread strides by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3266
- [NVRTC] Fix warp reduction compilation by @sepcnt in https://github.com/tile-ai/tilelang/pull/3260
- [BugFix][CUDA] Lower a kernel-body assert to a device-legal check by @xy200303 in https://github.com/tile-ai/tilelang/pull/3206
- [BugFix][JIT] Allocate a dynamic-shape output that precedes its sizing input by @xy200303 in https://github.com/tile-ai/tilelang/pull/3207
- [CUDA] Fuse exact FP4 to FP8 conversion through FP32 by @ZenAlexa in https://github.com/tile-ai/tilelang/pull/3204
- [Fix]Clamp T.copy/T.async_copy coalesced_width to achievable vector size instead of LOG(FATAL) by @edragain2nd in https://github.com/tile-ai/tilelang/pull/3246
- [BugFix] Bind loop targets independently of mutable scalar variables by @sepcnt in https://github.com/tile-ai/tilelang/pull/3232
- [BugFix][Language] Lower NaN-propagating clamp through device templates by @xy200303 in https://github.com/tile-ai/tilelang/pull/3205
- [Fix] Reject unsupported reduction NaN propagation dtypes by @ZenAlexa in https://github.com/tile-ai/tilelang/pull/3273
- [BugFix] Fix parallel loop lowering with let inlining disabled by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3269
- [Transform][CUDA] Fix async copy lowering with partitioned layouts by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3278
- [CUDA] Keep FP8 vector copies packed by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3276
- [CUDA] Support SM120 block-scaled GEMM fragments and odd warp atom grids by @sepcnt in https://github.com/tile-ai/tilelang/pull/3257
- [CUDA] Reject conflicting SM120 scale fragment layouts by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3284
- [Build] Optimize default single-config builds by @KellyFrog in https://github.com/tile-ai/tilelang/pull/3191
- [Fix][TVM] Update TVM to preserve dynamic reduction tail guards by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3294
- [BugFix] Eliminate unused bindings in Simplify by @LJC00118 in https://github.com/tile-ai/tilelang/pull/3293
- [Build][Windows] Restore ccache hits for clang-cl wheel builds by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3305
- [Public Release 9/30] Introduce Ascend 950 backend by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3308
- [Release] Bump version into 0.1.15 by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3309
New Contributors
- @Dino1844 made their first contribution in https://github.com/tile-ai/tilelang/pull/2917
- @mocusez made their first contribution in https://github.com/tile-ai/tilelang/pull/3145
- @rishabhsinha17 made their first contribution in https://github.com/tile-ai/tilelang/pull/3142
- @Guan-jeans made their first contribution in https://github.com/tile-ai/tilelang/pull/3126
- @anerli made their first contribution in https://github.com/tile-ai/tilelang/pull/3201
- @xuebozhang525-alt made their first contribution in https://github.com/tile-ai/tilelang/pull/3224
- @xy200303 made their first contribution in https://github.com/tile-ai/tilelang/pull/3206
Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.14...v0.1.15