| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| tilelang-0.1.14.tar.gz | 2026-09-02 | 91.7 MB | |
| tilelang-0.1.14-cp39-abi3-win_amd64.whl | 2026-09-02 | 28.5 MB | |
| tilelang-0.1.14-cp39-abi3-manylinux_2_34_aarch64.whl | 2026-09-02 | 41.0 MB | |
| tilelang-0.1.14-cp39-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl | 2026-09-02 | 45.4 MB | |
| tilelang-0.1.14-cp39-abi3-macosx_11_0_arm64.whl | 2026-09-02 | 32.7 MB | |
| README.md | 2026-09-01 | 24.8 kB | |
| v0.1.14 source code.tar.gz | 2026-09-01 | 12.6 MB | |
| v0.1.14 source code.zip | 2026-09-01 | 13.6 MB | |
| Totals: 8 Items | 265.5 MB | 0 | |
Highlights
- Reducer v2 (#2940, [#3093], [#3043], [#3044], [#3079], [#3100]):
T.alloc_reducerreworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates. - Warp specialization schedules (#2892): new scheduling and materialization mechanism for warp-specialized kernels.
- Layout inference cost models (#2960, [#3055], [#3061]): new IO-aware cost model for free-mode layout selection; register-count restored as the default, with an environment override to switch models.
- Unified backend resolution policy (#2318) plus backend split-up (#2855, [#2870], [#2850]): backend selection is now resolved through a single policy, and builtin ops / Python op proxies are split per backend (CUDA/ROCm/Metal).
- Compilation speed: up to ~4x faster cold parallel/AOT compilation (#2809); Z3 solvers materialized lazily (#3105) and analyzer contexts isolated per kernel compilation (#2890).
- TMA rework: TMA copy lowering unified on CuTe algebra (#3106); TMA layouts made region-aware to keep slices contiguous (#3089).
Language
- Recycle
T.unroll(explicit=True)for early explicit unrolling (#2859) - Make the region bridge a builtin intrinsic (#2983)
- Expose
cluster_maskonT.tma_copy(#2932) - Unify contiguous stride construction under a single implementation (#3016); honor declared strides in pointer helpers (#3073)
- Stricter validation: reject symbolic
T.gemmtile dimensions with a clear message (#3113), validateT.gemmk_packarguments (#3094), reject non-positivearrive_countinalloc_barrier/alloc_cluster_barrier(#3112), rejectT.Parallelindexing of local buffers (#3041), rejectbreakin fully expanded loops (#3078)
CUDA
- tcgen05: pack logical TMEM buffers into shared
tcgen05.allocarenas (#2831); support half-subpartition (M=64) TMEM tiles intcgen05.ld/st(#2880); fix ld/st segment pointer advancement in b32 columns (#2952) - Select the widest legal WGMMA N instead of gcd (#2931)
- FP32x2 accumulation for reductions: per-reduce control (#3057) and a global PassConfig (#3128)
- Pre-SM80 fallback for bf16 atomic add (#2938);
int4x2/uint4x2codegen (#3036); 16-bit CUTLASS type overloads for fast-math,__ldg, and htan intrinsics (#3097, [#3077], [#3028], [#2894]) - Fixes: warp shuffle for half/bfloat16/FP8 (#3056), FP8 min/max codegen (#3047), UB in packed 8-bit vector stores (#3092), logical not for vectorized bool (#3117, also HIP), vectorized Select codegen (#2843), ldmatrix source offsets wrapped within shared-memory regions (#3110), NVRTC kernel handles isolated per adapter (#2950), flat CUDA include discovery for NVRTC (#2829), masked warpsync in in-warp allreduce (#2865)
- Remove
T.{reads,writes}forT.tma_{gather4,scatter4}(#3053); separate TMA atomic-add dtype support from layout encoding (#2846)
ROCm and other backends
- Remove the Composable Kernel dependency (#3111); ROCm CI re-enabled on a gfx942 runner (#2874, [#2910])
- Fixes: preserve FP8 bits in warp shuffles (#3104), lower vector Select conditions lane-wise (#2889), emit a compiler barrier for
tl.sync_warpon HIP (#2872), reject sub-wavefront block sizes instead of crashing (#2918), resolve versioned device properties in the HIP stub (#2919); emit#linedirectives for the HIP target (#3058) - CPU backend: support atomic ops (#2941) and reduce ops (#2893)
- Metal: preserve pointer address spaces for byte offsets (#2925); resolve auto backend to torch and skip disk cache for torch (#2856)
- CuTeDSL: port backend intrinsics to CUTLASS DSL primitives (#2871)
Compiler / Transform
- Refactor the loop vectorization plan with ConstraintKind (#2935); always vectorize
T.Parallelloops (#3121); scalarize Select in automatic vectorization (#3060) - Add
VerifyBufferInit, a general buffer-initialization check (#2956) - Debug info: preserve source spans across lowering passes (#2966); emit
#linedirectives from TIR spans (#3048) - Fixes: don't drop syncs from the other if branch (#3085), fix wait parity for explicit mbarriers in pipelined loops (#3087), avoid int32 overflow in vector analysis (#3066), fix non-divisible nested modulo simplification (#3065), fix ties-away-from-zero round compile on bfloat16/float8 (#2873), fix absmax/abssum for uint dtypes (#2845), carry memory_order through vectorized atomic_add (#2924), fix unsigned zero-point decode underflow (#3118), keep cp.async operands in their address spaces (#2869), handle grid barriers and unbounded pointer ranges (#3050), deduplicate replicated reducer updates (#2881), bind symbolic coordinate ranges in FragmentThreadIndexProbe (#3096), reject non-round-tripping inferred layout inverses (#3090)
- Cleanup: remove obsolete compiler and runtime paths (#3086); remove the obsolete disable-fast-math pass config (#3098)
Runtime / JIT / Build
- Kernel cache: detect and repair corrupted cache entries (#3074), export libraries after disk-cache hits (#3116), remove the separate cache temporary directory (#3069)
- Allocate kernel outputs through the packed API (#2937); fix
get_parent_localsframe self-reference leak (#2934) - Support Cython 3.3 with the Python 3.9 limited API (#3068); fix CMake reconfigure aborting in FindPipCUDAToolkit before
project()(#3102)
Tooling / Ecosystem
- Official compile-only CLI (#3045)
- Unified pass instrumentation per compilation (#2923); Pass Visualizer driven by PassInstrument (#2866); show kernel name in the pass timing report (#2905)
- Data race check disabled by default, opt-in via env var (#2851)
- Open-source TileLang LSP announced (#2862); new agent skills: simplification (#3080), semantic validation (#3054), PR submission (#3082), backend architecture (#2900)
- Examples: generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling (#2902); adopt multi-staged buffers in examples (#2836)
- Docs: add Sunrise-AI TANG (#3107), HYGON (#3101), and MetaX MACA (#3095) to supported platforms; link the multi-backend architecture design (#3123)
What's Changed
- [CUDA] Pack logical TMEM buffers into shared
tcgen05.allocarenas by @Rachmanino in https://github.com/tile-ai/tilelang/pull/2831 - [Enhancement] Speed up cold parallel/AOT compilation up to ~4x by @cklxx in https://github.com/tile-ai/tilelang/pull/2809
- [Docs] Refresh README news and onboarding by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2849
- [Enhancement] Disable data race check by default, opt-in via env var by @KellyFrog in https://github.com/tile-ai/tilelang/pull/2851
- [BugFix] Fix absmax and abssum for uint dtypes by @jjppp in https://github.com/tile-ai/tilelang/pull/2845
- [CUDA][TMA] Separate atomic-add dtype support from layout encoding by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2846
- [TIR][Python] Trim redundant op proxy wrappers by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2850
- [CUDA][ROCm][Metal] Split backend-specific builtin ops by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2855
- [CUDA] Adopt multi-staged buffers in examples by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/2836
- [CI] [pre-commit.ci] autoupdate by @pre-commit-ci[bot] in https://github.com/tile-ai/tilelang/pull/2861
- [Docs][LSP] Announce the open-source TileLang LSP by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2862
- [BugFix][Metal] Resolve auto backend to torch and skip disk cache for torch by @oraluben in https://github.com/tile-ai/tilelang/pull/2856
- [Typo] Correct source spelling errors by @morluto in https://github.com/tile-ai/tilelang/pull/2858
- [Refactor] Recycle
T.unroll(explicit=True)for early explicit unrolling by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/2859 - [BugFix] Use masked warpsync in in-warp allreduce by @jjppp in https://github.com/tile-ai/tilelang/pull/2865
- [Debug][TIR] Drive Pass Visualizer with PassInstrument by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2866
- [Doc] Fix wrong loop bound in FlashAttention README example by @shanyi0228-web in https://github.com/tile-ai/tilelang/pull/2868
- [Examples] Gate CUDA-only and flash_attn-dependent example tests by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2864
- [Testing] Gate CUDA-only tests so non-CUDA backends can run the suite by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2863
- [TIR][Python] Split backend-specific op proxies by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2870
- [BugFix][ROCm] Emit a compiler barrier for tl.sync_warp on HIP by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2872
- [BugFix] Handle vectorized SelectNode in codegen_cuda by @jjppp in https://github.com/tile-ai/tilelang/pull/2843
- [CI] Re-enable ROCm CI on a gfx942 runner by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2874
- [Cleanup] Replace root reproducers with CPU regression coverage by @GY-Bai in https://github.com/tile-ai/tilelang/pull/2878
- [Compiler][Z3] Isolate analyzer contexts per kernel compilation by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2890
- [Backend] Add unified backend resolution policy by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2318
- [Doc] ROCm CI is no longer disabled by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2896
- [BugFix] Deduplicate replicated reducer updates by @KellyFrog in https://github.com/tile-ai/tilelang/pull/2881
- [Test] Add regression test for issue [#2883] by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2899
- [CPU] Support reduce ops on CPU by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/2893
- [Docs] Define backend architecture and integration skill by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2900
- [Testing] Gate the issue [#2883] regression test on CUDA by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2908
- [BugFix][ROCm] Lower vector Select conditions lane-wise by @morluto in https://github.com/tile-ai/tilelang/pull/2889
- [CI] Use stable torch for the ROCm leg by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2910
- [Testing] Run portable regression tests on auto targets by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2914
- [BugFix][Carver] Parse lettered SM arch strings in check_sm_version by @adityasingh2400 in https://github.com/tile-ai/tilelang/pull/2891
- [Enhancement] Show kernel name in pass timing report by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/2905
- [CuTeDSL] Port backend intrinsics to CUTLASS DSL primitives by @cherichy in https://github.com/tile-ai/tilelang/pull/2871
- [BugFix][ROCm] Reject sub-wavefront block sizes instead of crashing by @andyluo7 in https://github.com/tile-ai/tilelang/pull/2918
- [Fix] Discover flat CUDA includes for NVRTC by @morluto in https://github.com/tile-ai/tilelang/pull/2829
- [CI]: Bump pypa/cibuildwheel from 4.1 to 4.2 by @dependabot[bot] in https://github.com/tile-ai/tilelang/pull/2930
- [ROCm] Resolve versioned device properties in HIP stub by @skyguan92 in https://github.com/tile-ai/tilelang/pull/2919
- [BugFix] Skip DecoupleTypeCast on Evaluate roots to keep cp.async operands in their address spaces by @li-ruinan in https://github.com/tile-ai/tilelang/pull/2869
- [Debug][TIR][JIT] Unify pass instrumentation per compilation by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2923
- [Cherry-Pick][BugFix] Fix get_parent_locals frame self-reference leak by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2934
- [BugFix] Fix failed ties-away-from-zero round compile on bfloat16/float8 by @edragain2nd in https://github.com/tile-ai/tilelang/pull/2873
- [JIT][FFI] Allocate kernel outputs through the packed API by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2937
- [Refactor][BugFix] Refactor the loop vectorization plan with ConstraintKind by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2935
- [BugFix][CUDA] Provide htan overloads for fp16/bf16 tangent by @Ruihan11 in https://github.com/tile-ai/tilelang/pull/2894
- [Layout] Support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st by @Rachmanino in https://github.com/tile-ai/tilelang/pull/2880
- [Feature] Warp specialization schedules and materialization by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/2892
- [CUDA] Add pre-SM80 fallback for bf16 atomic add by @Chennesxu in https://github.com/tile-ai/tilelang/pull/2938
- [BugFix][Hopper] Select the widest legal WGMMA N instead of gcd by @bigSheep123 in https://github.com/tile-ai/tilelang/pull/2931
- [Enhancement] Expose cluster_mask on T.tma_copy by @bigSheep123 in https://github.com/tile-ai/tilelang/pull/2932
- [CPU] Support atomic ops on CPU by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/2941
- [Example] Generalize TCGEN05 GEMMs for Thor (sm110a) and add Stream-K scheduling by @xinhao-luo in https://github.com/tile-ai/tilelang/pull/2902
- [Layout][CUDA] Support reinterpreting (dtype-changing) T.view aliases by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/2953
- [BugFix][CUDA] Advance tcgen05 ld/st segment pointers in b32 columns by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/2952
- [Testing] Pin the folded-base descriptor form for static ts slices by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/2951
- [Lang] Reducer v2: first-class deferred reduction epochs with planned physical lowering by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2940
- [CI][Examples] Remove TopK example from performance regression by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2958
- [BugFix][Metal] Preserve pointer address spaces for byte offsets by @GY-Bai in https://github.com/tile-ai/tilelang/pull/2925
- [BugFix] Isolate NVRTC kernel handles per adapter by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/2950
- [Lang][TIR] Make region bridge a builtin intrinsic by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2983
- [Layout][Inference] Add IO-aware cost model for free-mode selection by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2960
- [Transform] Add VerifyBufferInit, a general buffer-initialization check by @RyanL2 in https://github.com/tile-ai/tilelang/pull/2956
- [CUDA] Add __ldg overloads for 16-bit CUTLASS types by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3028
- [Analysis] Reject
T.Parallelindexing of local buffers by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3041 - [Lang][Reducer] Support loop-scoped epochs and legacy default allocations by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3043
- [TIR][Transform] Preserve source spans across lowering passes by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/2966
- [Lang][Reducer] Allow conditional reducer finalization by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3044
- [CUDA] Fix FP8 min/max codegen by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3047
- [BugFix][Layout] Avoid thread-indexed wide reducer finalize readback by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3049
- [TIR][Transform] Handle grid barriers and unbounded pointer ranges by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3050
- [CodeGen] Emit #line directives from TIR spans by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3048
- [Misc] Remove incorrect ASF license headers from src files by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3051
- [Bugfix] Carry memory_order through vectorized atomic_add by @arcusbuilds in https://github.com/tile-ai/tilelang/pull/2924
- [Layout][Inference] Restore register-count as the default cost model by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3055
- [CUDA] Remove
T.{reads,writes}forT.tma_{gather4,scatter4}by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3053 - [Skill] Add TileLang semantic validation skill by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3054
- [CUDA][Reduce] Add per-reduce control for FP32x2 accumulation by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3057
- [Layout][Config] Add environment override for layout cost model by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3061
- [CUDA] Fix warp shuffle for half, bfloat16, and FP8 by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3056
- [Build] Support Cython 3.3 with the Python 3.9 limited API by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3068
- [Runtime][Cache] Remove separate cache temporary directory by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3069
- [Language] Honor declared strides in pointer helpers by @zupengwang in https://github.com/tile-ai/tilelang/pull/3073
- [BugFix] Scalarize Select in automatic vectorization by @KellyFrog in https://github.com/tile-ai/tilelang/pull/3060
- [CodeGen][ROCm] Emit #line directives for HIP target by @penguin-wwy in https://github.com/tile-ai/tilelang/pull/3058
- [Runtime][Cache] Detect and repair corrupted cache entries by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3074
- [Tool] Add official compile-only CLI by @LibertychaserUS in https://github.com/tile-ai/tilelang/pull/3045
- [BugFix][CUDA] Support int4x2 and uint4x2 codegen by @SamJSui in https://github.com/tile-ai/tilelang/pull/3036
- [CUDA] Bridge half-style math intrinsics for 16-bit CUTLASS types by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3077
- [Bugfix] Fix non-divisible nested modulo simplification by @haoyang9804 in https://github.com/tile-ai/tilelang/pull/3065
- [Docs][CI] Add TileLang PR submission skill by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3082
- [Lang][Reducer] Vectorize contiguous reducer updates by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3079
- [SKILL] Add TileLang simplification skill by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3080
- [Refactor] Remove obsolete compiler and runtime paths by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3086
- [Language] Unify contiguous stride construction using a single implem… by @jjppp in https://github.com/tile-ai/tilelang/pull/3016
- [Bugfix] Fix wait parity for explicit mbarriers in pipelined loops by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3087
- [Bugfix] Make TMA layouts region-aware to keep slices contiguous by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3089
- [Docs] Fix stale repository links by @morluto in https://github.com/tile-ai/tilelang/pull/2857
- [Refactor] Give reducers a first-class PartialFragment layout solved by layout inference by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3093
- [BugFix] Bind symbolic coordinate ranges in FragmentThreadIndexProbe by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3096
- [Doc] Update support info with MetaX MACA backend by @Five-HZ in https://github.com/tile-ai/tilelang/pull/3095
- [CUDA] Add 16-bit overloads for CUTLASS fast-math functions by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3097
- [Layout] Reject non-round-tripping inferred inverses by @KellyFrog in https://github.com/tile-ai/tilelang/pull/3090
- [Bugfix] Validate T.gemm k_pack arguments by @WenzheWang in https://github.com/tile-ai/tilelang/pull/3094
- [Cleanup] Remove obsolete disable-fast-math pass config by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3098
- [Bugfix][CUDA] Fix UB in packed 8-bit vector stores by @yydhYYDH in https://github.com/tile-ai/tilelang/pull/3092
- [BugFix][Transform] Avoid int32 overflow in vector analysis by @kobecai in https://github.com/tile-ai/tilelang/pull/3066
- [Layout][Reducer] Preserve vectorized reducer update plans by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3100
- [Docs] List HYGON in README ecosystem hardware adapters by @warrenzzhou in https://github.com/tile-ai/tilelang/pull/3101
- [Docs] Add Sunrise-AI TANG to platform support by @cratoroo in https://github.com/tile-ai/tilelang/pull/3107
- [Bugfix][CMake] Fix reconfigure aborting in FindPipCUDAToolkit before project() by @SuperGoodGame in https://github.com/tile-ai/tilelang/pull/3102
- [Compiler][Z3] Bump 3rdparty/tvm to materialize Z3 solvers lazily by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3105
- [ROCm] Preserve FP8 bits in warp shuffles by @andyluo7 in https://github.com/tile-ai/tilelang/pull/3104
- [CUDA] Unify TMA copy lowering on CuTe algebra by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3106
- [ROCm] Remove the Composable Kernel dependency by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3111
- [Lang] Reject non-positive arrive_count in alloc_barrier and alloc_cluster_barrier by @yurekami in https://github.com/tile-ai/tilelang/pull/3112
- [Ci] Register TVM feature markers to get rid of PytestUnknownMarkWarning by @jjppp in https://github.com/tile-ai/tilelang/pull/3115
- [Bugfix][Quantize] Fix unsigned zero-point decode underflow (#2947) by @Junius-Wynn in https://github.com/tile-ai/tilelang/pull/3118
- [Bugfix] Reject symbolic T.gemm tile dimensions with a clear message by @jjppp in https://github.com/tile-ai/tilelang/pull/3113
- [BugFix][CUDA][HIP] Handle logical not for vectorized bool by @jjppp in https://github.com/tile-ai/tilelang/pull/3117
- [CUDA] Wrap ldmatrix source offsets within shared-memory regions by @Chennesxu in https://github.com/tile-ai/tilelang/pull/3110
- [Runtime][Cache] Export libraries after disk-cache hits by @ZenAlexa in https://github.com/tile-ai/tilelang/pull/3116
- [BugFix][Transform] Reject break in fully expanded loops (#3026) by @KellyFrog in https://github.com/tile-ai/tilelang/pull/3078
- [BugFix] Always vectorize
T.Parallelloops by @Yongqi-Zhuo in https://github.com/tile-ai/tilelang/pull/3121 - [BugFix] Don't drop syncs from the other if branch by @haoyang9804 in https://github.com/tile-ai/tilelang/pull/3085
- [Docs] Link multi-backend architecture design by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3123
- [Release] Bump version to 0.1.14 by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/3119
- [CUDA][Reduce] Add PassConfig for FP32x2 accumulation by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/3128
New Contributors
- @shanyi0228-web made their first contribution in https://github.com/tile-ai/tilelang/pull/2868
- @andyluo7 made their first contribution in https://github.com/tile-ai/tilelang/pull/2864
- @adityasingh2400 made their first contribution in https://github.com/tile-ai/tilelang/pull/2891
- @skyguan92 made their first contribution in https://github.com/tile-ai/tilelang/pull/2919
- @edragain2nd made their first contribution in https://github.com/tile-ai/tilelang/pull/2873
- @Ruihan11 made their first contribution in https://github.com/tile-ai/tilelang/pull/2894
- @bigSheep123 made their first contribution in https://github.com/tile-ai/tilelang/pull/2931
- @xinhao-luo made their first contribution in https://github.com/tile-ai/tilelang/pull/2902
- @RyanL2 made their first contribution in https://github.com/tile-ai/tilelang/pull/2956
- @arcusbuilds made their first contribution in https://github.com/tile-ai/tilelang/pull/2924
- @zupengwang made their first contribution in https://github.com/tile-ai/tilelang/pull/3073
- @LibertychaserUS made their first contribution in https://github.com/tile-ai/tilelang/pull/3045
- @SamJSui made their first contribution in https://github.com/tile-ai/tilelang/pull/3036
- @haoyang9804 made their first contribution in https://github.com/tile-ai/tilelang/pull/3065
- @Five-HZ made their first contribution in https://github.com/tile-ai/tilelang/pull/3095
- @WenzheWang made their first contribution in https://github.com/tile-ai/tilelang/pull/3094
- @yydhYYDH made their first contribution in https://github.com/tile-ai/tilelang/pull/3092
- @kobecai made their first contribution in https://github.com/tile-ai/tilelang/pull/3066
- @warrenzzhou made their first contribution in https://github.com/tile-ai/tilelang/pull/3101
- @cratoroo made their first contribution in https://github.com/tile-ai/tilelang/pull/3107
- @SuperGoodGame made their first contribution in https://github.com/tile-ai/tilelang/pull/3102
- @Junius-Wynn made their first contribution in https://github.com/tile-ai/tilelang/pull/3118
- @ZenAlexa made their first contribution in https://github.com/tile-ai/tilelang/pull/3116
Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.13...v0.1.14