| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| transformer_engine_torch-2.18.0+cu13torch26.06cxx11abiTRUE-cp312-cp312-linux_x86_64.whl | 2026-08-14 | 1.1 MB | |
| transformer_engine_torch-2.18.0+cu13torch26.07cxx11abiTRUE-cp312-cp312-linux_x86_64.whl | 2026-08-14 | 1.1 MB | |
| transformer_engine_torch-2.18.0+cu13torch26.05cxx11abiTRUE-cp312-cp312-linux_x86_64.whl | 2026-08-14 | 1.1 MB | |
| transformer_engine_torch-2.18.0+cu13torch26.04cxx11abiTRUE-cp312-cp312-linux_x86_64.whl | 2026-08-14 | 1.1 MB | |
| README.md | 2026-08-14 | 5.9 kB | |
| v2.18 source code.tar.gz | 2026-08-14 | 5.7 MB | |
| v2.18 source code.zip | 2026-08-14 | 6.4 MB | |
| Totals: 7 Items | 16.3 MB | 0 | |
Transformer Engine v2.18 Release Notes
Key Features and Enhancements
- [Common] Added fused BF16 forward/backward kernels for scaled variants of
SwiGLU,ClampedSwiGLU, andSReLU. (#3132) - [Common] Reduced binary size and build time by migrating fused softmax, LayerNorm, and RMSNorm kernels to NVRTC. (#3156)
- [Common] Improved the performance of variable-length cuDNN fused attention forward passes by passing sequence-length metadata directly to cuDNN. (#3186)
- [Common, PyTorch] Added dense
topk_indicesoutput support to the fused router for TE Expert Parallelism. (#3129) - [Common, PyTorch] Added FP8 current-scaling grouped quantization with support for varying tensor dimensions. (#3114)
- [Common, PyTorch] Made FP8 current scaling in
GroupedLinearcompatible with CUDA graph capture and improved cuBLASLt grouped GEMM algorithm selection. (#3143) - [Common, PyTorch] Added CUDA graph capture support for THD attention. (#2898)
- [Common, JAX, PyTorch] Enabled cuDNN fused attention backward for attention head dimension 256 on Blackwell GPUs with compute capability 10.x. (#3056)
- [PyTorch] Enabled FP8 block scaling in
GroupedLinearand fusible ops. (#3171) (#3242) (#3135) - [PyTorch] Reduced scale-factor swizzling overhead for cached MXFP8 and NVFP4 weights. (#3093)
- [PyTorch] Added THD format support for AllGather-based context parallelism. (#2829)
- [PyTorch] Added zero-copy support to TE Expert Parallelism. (#3187)
- [PyTorch] Enabled NVFP4 RHT amax fusion in the grouped MLP using SRelu activation. (#3133)
- [PyTorch] Added
NVTE_FLASH_ATTN_V2,NVTE_FLASH_ATTN_V3, andNVTE_FLASH_ATTN_V4environment variables to control each FlashAttention version independently. (#3204) - [JAX] Integrated the TE EP backend into the experimental
moe()VJP API and experimental Flax MoEBlock. (#3116) - [Docs] Improved the documentation on attention (#3142) (#3162)
- [Build] Replaced the cuDNN Frontend submodule with the
nvidia-cudnn-frontendpython package for source builds. (#3169) (#3278)
Fixed Issues
- [Common, PyTorch] Fixed grouped GEMM failures with FP8 per-tensor scaling on Hopper when using cuBLAS versions earlier than 13.5 by falling back to the legacy path. (#3181)
- [Common, PyTorch] Fixed CUDA graph failures with cached NVFP4 weights by producing GEMM-swizzled scaling factors during quantization when supported and disabling pre-swizzling otherwise. (#3190) (#3232)
- [PyTorch] Fixed gradient propagation for quantized tensors across autograd boundaries and made FP8 block-scaling dequantization autograd-aware. (#3172)
- [PyTorch] Fixed
NVTE_BACKWARD_OVERRIDE=dequantizedto use dequantized forward operands in backward. (#3141) - [PyTorch] Fixed precision loss in the fused cross-entropy backward pass. (#3193)
- [PyTorch] Fixed process aborts during communication+GEMM overlap and NCCL-EP initialization caused by incorrect GIL handling. (#3203)
- [PyTorch] Fixed skipped updates in
FusedAdamandFusedSGDwhen parameter groups contain empty tensors. (#3212) - [PyTorch] Fixed NCCL communicator initialization when creating a cuSOLVERMp context. (#3240)
- [PyTorch] Fixed quantized tensor
__repr__failures when data cannot be materialized, including undertorch.compile. (#3146) - [JAX] Fixed the distributed MoE auxiliary loss backward pass by preserving the scalar cotangent. (#3237)
- [Build] Fixed a CUDA 13.0 source build failure in grouped FP8 quantization code. (#3194)
Known Issues in This Release
- There is a known compatibity issue between the FlashAttention v4, CuTeDSL and CUDNN Frontend pip packages, which could produce runtime errors like
ImportError: cannot import name 'block_copy' from 'cutlass.utils'. A supported stable combination: flash-attn-4==4.0.0b11nvidia-cutlass-dsl[cu13]==4.4.2nvidia-cudnn-frontend==1.26.0
Breaking Changes in This Release
There are no breaking changes in this release.
Deprecated Features
There are no deprecated features in this release.