Download Latest Version libtorchtrt-2.14.0-tensorrt11.1.0-cuda132-libtorch2.14.0-x86_64-windows.zip (956.5 kB) Google Add to Preferred Sources
Home / v2.14.0
Name Modified Size InfoDownloads / Week
Parent folder
torch_tensorrt-2.14.0-cp310-cp310-win_amd64.whl 2026-09-23 2.2 MB
torch_tensorrt-2.14.0-cp312-cp312-win_amd64.whl 2026-09-23 2.2 MB
torch_tensorrt-2.14.0-cp311-cp311-win_amd64.whl 2026-09-23 2.2 MB
torch_tensorrt-2.14.0-cp313-cp313-win_amd64.whl 2026-09-23 2.2 MB
torch_tensorrt-2.14.0-cp314-cp314-win_amd64.whl 2026-09-23 2.2 MB
torch_tensorrt-2.14.0-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-23 4.3 MB
torch_tensorrt-2.14.0-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-23 4.3 MB
torch_tensorrt-2.14.0-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-23 4.3 MB
torch_tensorrt-2.14.0-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-23 4.3 MB
torch_tensorrt-2.14.0-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-23 4.3 MB
libtorchtrt-2.14.0-tensorrt11.1.0-cuda132-libtorch2.14.0-aarch64-linux.tar.gz 2026-09-23 3.0 MB
libtorchtrt-2.14.0-tensorrt11.1.0-cuda132-libtorch2.14.0-x86_64-linux.tar.gz 2026-09-23 3.2 MB
torch_tensorrt-2.14.0-cp310-cp310-manylinux_2_28_aarch64.whl 2026-09-23 4.2 MB
torch_tensorrt-2.14.0-cp311-cp311-manylinux_2_28_aarch64.whl 2026-09-23 4.2 MB
torch_tensorrt-2.14.0-cp312-cp312-manylinux_2_28_aarch64.whl 2026-09-23 4.2 MB
torch_tensorrt-2.14.0-cp313-cp313-manylinux_2_28_aarch64.whl 2026-09-23 4.2 MB
torch_tensorrt-2.14.0-cp314-cp314-manylinux_2_28_aarch64.whl 2026-09-23 4.2 MB
torch_tensorrt_executorch_runtime-0.1.0-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-23 7.6 MB
torch_tensorrt_executorch_runtime-0.1.0-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-23 7.6 MB
torch_tensorrt_executorch_runtime-0.1.0-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-23 7.6 MB
torch_tensorrt_executorch_runtime-0.1.0-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-23 7.6 MB
torch_tensorrt_executorch_runtime-0.1.0-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-23 7.6 MB
torch_tensorrt_executorch_runtime-0.1.0-cp314-cp314-manylinux_2_28_aarch64.whl 2026-09-23 7.3 MB
torch_tensorrt_executorch_runtime-0.1.0-cp313-cp313-manylinux_2_28_aarch64.whl 2026-09-23 7.3 MB
torch_tensorrt_executorch_runtime-0.1.0-cp312-cp312-manylinux_2_28_aarch64.whl 2026-09-23 7.3 MB
torch_tensorrt_executorch_runtime-0.1.0-cp311-cp311-manylinux_2_28_aarch64.whl 2026-09-23 7.3 MB
torch_tensorrt_executorch_runtime-0.1.0-cp310-cp310-manylinux_2_28_aarch64.whl 2026-09-23 7.3 MB
torch_tensorrt_rtx-2.14.0-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-23 3.9 MB
torch_tensorrt_rtx-2.14.0-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-23 3.9 MB
torch_tensorrt_rtx-2.14.0-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-23 3.9 MB
torch_tensorrt_rtx-2.14.0-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-23 3.9 MB
torch_tensorrt_rtx-2.14.0-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-23 3.9 MB
libtorchtrt-2.14.0-tensorrt11.1.0-cuda132-libtorch2.14.0-x86_64-windows.zip 2026-09-23 956.5 kB
torch_tensorrt_rtx-2.14.0-cp313-cp313-win_amd64.whl 2026-09-23 2.1 MB
torch_tensorrt_rtx-2.14.0-cp314-cp314-win_amd64.whl 2026-09-23 2.1 MB
torch_tensorrt_rtx-2.14.0-cp312-cp312-win_amd64.whl 2026-09-23 2.1 MB
torch_tensorrt_rtx-2.14.0-cp311-cp311-win_amd64.whl 2026-09-23 2.1 MB
torch_tensorrt_rtx-2.14.0-cp310-cp310-win_amd64.whl 2026-09-23 2.1 MB
torch_tensorrt_rtx-2.14.0+cu134-cp313-cp313-win_arm64.whl 2026-09-10 2.0 MB
README.md 2026-09-09 24.1 kB
Torch-TensorRT v2.14.0 source code.tar.gz 2026-09-09 552.7 MB
Torch-TensorRT v2.14.0 source code.zip 2026-09-09 608.7 MB
Totals: 42 Items   1.3 GB 4

Torch-TensorRT 2.14.0 Linux x86-64 and Windows targets

PyTorch 2.14, CUDA 12.6/13.0/13.2, TensorRT 11.1, Python 3.10~3.14

Torch-TensorRT Wheels are available:

x86-64 Linux and Windows: CUDA 13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

CUDA 12.6/13.0/13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

aarch64 SBSA Linux and Jetson Thor and Orin: CUDA 13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

CUDA 13.0/13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

Torch-TensorRT-RTX 2.14.0 Linux x86-64 and Windows targets

PyTorch 2.14, CUDA 12.6/13.0/13.2, TensorRT-RTX 1.6, Python 3.10~3.14

Torch-TensorRT-RTX Wheels are available:

x86-64 Linux and Windows: CUDA 13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

CUDA 12.6/13.0/13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

Torch-TensorRT-Executorch-Runtime 2.14.0 Linux x86-64 and Windows targets

PyTorch 2.14, CUDA 12.6/13.0/13.2, TensorRT 11.1, Python 3.10~3.14

Torch-TensorRT-Executorch-Runtime Wheels are available:

x86-64 Linux: CUDA 13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

CUDA 12.6/13.0/13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

aarch64 SBSA Linux and Jetson Thor and Orin: CUDA 13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

CUDA 12.6/13.0/13.2 + Python 3.10-3.14 + Torch 2.14 + TensorRT 11.1

New Features:

Executorch

Torch-TensorRT 2.14 expands its ExecuTorch integration with a more complete export and deployment workflow for TensorRT-accelerated .pte models.

Highlights:

  • New torch_tensorrt.executorch.export() API for producing an inspectable EdgeProgramManager before serialization.
  • Support for exporting multiple named methods into a single .pte artifact, with optional constant methods and ETRecord generation for debugging.
  • torch_tensorrt.save(..., output_format="executorch") remains the streamlined path for one-step model export.
  • Supports composing the TensorRT partitioner with additional ExecuTorch partitioners, enabling mixed TensorRT and CUDA delegation in one artifact.
  • The optional torch-tensorrt-executorch-runtime package enables loading and running exported .pte files through torch_tensorrt.load(..., format="executorch").
  • Continues to use the lightweight native TensorRT ExecuTorch backend, without a Torch or LibTorch dependency at deployment time.
  • Aliased I/O for in-place operations — TensorRT engines can now write directly into a caller-owned or module-owned tensor's storage, so KV-cache updates in streaming and autoregressive inference no longer pay a full cache-sized copy at the engine boundary.
  • Multiple optimization profiles — A single engine can now carry several independently tuned shape profiles selected at runtime by index or automatically, so bimodal workloads like LLM prefill and decode each get kernels tuned for their own shapes instead of splitting the difference.

ExecuTorch support is currently available on Linux(x86-64+aarch64), will Support Windows from 2.15 release.

Multi-Device in TRT-RTX

This release expands multi-GPU support across standard TensorRT and TensorRT-RTX configurations. Native TensorRT collective operations are enabled for TRT-RTX 1.5+ builds on supported Ampere or newer GPUs. Also there is improved process group discovery, including non contiguous subgroup identifiers and serialized save/load, which prevents crashes and ensures distributed engines bind to the correct world group or subgroup.

FP8 Fused Attention Layers / Quantization

This release introduces native FP8 support for TensorRT's IAttention layer. To enable this and improve overall attention layer performance, four attention converters were comprehensively refactored to utilize TensorRT's updated add_attention_v2() API. Beyond core attention support, the update strengthens the FP8 quantization pipeline by adding support for ITensor amax computations, complete with a new export_torch_mode() wrapper to streamline the model export process. Finally, to ensure a smoother out-of-the-box developer experience, this update patches NVIDIA ModelOpt versioning issues within the provided Vision Transformer (ViT) FP8 quantization examples.

Aliased IO and KV Caching

Aliased I/O for in-place operations (#4251) TensorRT engines can now write in place into a tensor you own, removing the full-tensor copy that in-place operators previously paid at the engine boundary. The main beneficiary is streaming and autoregressive inference with a key/value cache, where each step wrote one timestep but copied the whole cache. Caches passed as inputs and caches registered as module buffers both work, on either runtime, with or without CUDA graphs, and compiled modules still save and load normally. Cases TensorRT can't alias fall back to the previous behavior, so nothing that compiled before stops compiling. This is an ABI-breaking change: previously serialized engines must be rebuilt.

Multi Optimization Profile

A single engine can now carry several independently tuned shape profiles and switch between them at runtime, so one set of weights serves multiple shape regimes without rebuilding. This matters most for autoregressive LLMs, where long prefill and single-token decode previously had to share one tuning point and decode ran on prefill-shaped kernels. Declare profiles with the new profiles argument to torch_tensorrt.Input, then pin one per with block using torch_tensorrt.runtime.optimization_profile, or pass "auto" to select by input shape. Engines that declare no profiles are unaffected.

Improved Complex number lowering in Dynamo

Torch-TensorRT now integrates Pytorch's upstream complex-number decomposition for dynamo compiled models. For torch versions >= 2.14, the new path is enabled automatically and converts complex operations into TensorRT compatible real valued operations while preserving complex I/O and symbolic dynamic shapes.

What's Changed

New Contributors

Full Changelog: https://github.com/pytorch/TensorRT/compare/v2.13.0...v2.14.0

Source: README.md, updated 2026-09-09