TileLang is a Pythonic domain-specific language for writing high-performance GPU, CPU, and accelerator kernels without dropping fully into low-level vendor code. Its compiler infrastructure is built on TVM and targets workloads such as GEMM, dequantization, FlashAttention, and linear attention. Developers can work at different abstraction levels while still controlling memory layout, scheduling, pipelines, and hardware-specific behavior. Current backends include CUDA, ROCm, Metal, CPU, and Huawei Ascend support. The project includes autotuning, debugging tools, layout visualization, compiler diagnostics, and an LSP for editor assistance. Kernels can be integrated with PyTorch and compiled for several modern accelerator families. TileLang aims to combine research-friendly productivity with performance close to specialized hand-tuned kernels.
Features
- Pythonic high-performance kernel language
- CUDA, ROCm, Metal, CPU, and Ascend backends
- GEMM, FlashAttention, and dequantization kernels
- Autotuning and hardware-aware scheduling
- LSP, diagnostics, and debugging tools
- PyTorch and TVM-based integration