Download Latest Version TorchRL 0.14_ DreamerV3, Async at Scale, and Composable Replay source code.zip (12.2 MB) Google Add to Preferred Sources
Home / v0.14.0
Name Modified Size InfoDownloads / Week
Parent folder
torchrl-0.14.0+cu130-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp312-cp312-manylinux_2_28_aarch64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp314-cp314-win_amd64.whl 2026-09-10 3.3 MB
torchrl-0.14.0+cu126-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0+cu130-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0+cu126-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp310-cp310-win_amd64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-10 3.4 MB
torchrl-0.14.0+cu126-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp313-cp313-win_amd64.whl 2026-09-10 3.3 MB
torchrl-0.14.0+cu132-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0+cu126-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0+cu130-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp314-cp314-macosx_14_0_arm64.whl 2026-09-10 3.7 MB
torchrl-0.14.0-cp312-cp312-macosx_14_0_arm64.whl 2026-09-10 3.7 MB
torchrl-0.14.0+cu132-cp312-cp312-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp313-cp313-macosx_14_0_arm64.whl 2026-09-10 3.7 MB
torchrl-0.14.0+cu126-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp310-cp310-macosx_14_0_arm64.whl 2026-09-10 3.7 MB
torchrl-0.14.0-cp314-cp314-manylinux_2_28_aarch64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-10 3.4 MB
torchrl-0.14.0+cu130-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp310-cp310-manylinux_2_28_aarch64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-10 3.4 MB
torchrl-0.14.0-cp311-cp311-win_amd64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp311-cp311-macosx_14_0_arm64.whl 2026-09-10 3.7 MB
torchrl-0.14.0+cu132-cp310-cp310-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp313-cp313-manylinux_2_28_aarch64.whl 2026-09-10 3.3 MB
torchrl-0.14.0+cu132-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-10 3.4 MB
torchrl-0.14.0+cu130-cp311-cp311-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
torchrl-0.14.0-cp312-cp312-win_amd64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp311-cp311-manylinux_2_28_aarch64.whl 2026-09-10 3.3 MB
torchrl-0.14.0-cp314-cp314-manylinux_2_28_x86_64.whl 2026-09-10 3.4 MB
torchrl-0.14.0+cu132-cp313-cp313-manylinux_2_28_x86_64.whl 2026-09-10 3.5 MB
README.md 2026-09-10 43.9 kB
TorchRL 0.14_ DreamerV3, Async at Scale, and Composable Replay source code.tar.gz 2026-09-10 11.1 MB
TorchRL 0.14_ DreamerV3, Async at Scale, and Composable Replay source code.zip 2026-09-10 12.2 MB
Totals: 38 Items   144.3 MB 0

TorchRL v0.14.0

TorchRL 0.14 is a major feature release spanning model-based reinforcement learning, distributed training and inference, replay-buffer composition, checkpointing, on-policy optimization, vision-language-action workflows, and simulation tooling. It also tightens several contracts announced in earlier releases and requires TensorDict 0.14.2 or later in the 0.14 release line.

Highlights

  • DreamerV3 receives an end-to-end correctness and performance overhaul, including a block-RSSM core, public DreamerV3BlockGRU modules, Triton and CUDA-graph acceleration, corrected losses and continuation semantics, reusable percentile return normalization, Hydra-compatible model configuration, and a reproducible DMC Walker benchmark. See #4072, #4074, #4075, #4125, #4126, #4132, #4133, #4158, #4160, #4174, #4187, #4196, and #4216 by @vmoens, @theap06, and @bsprenger, with correctness fixes in #4119 by @gtnv and #4118 and #4128 by @unography.
  • Distributed execution now covers configurable TorchRL service transports, shared-memory inference, process inference servers, DDP and Ray trainer execution, remote learners, and unified collector backend construction. See #3955, #3958, #4022, and the distributed trainer commit stack by @vmoens.
  • Replay buffers gain a trajectory query language, composable sample units, rich sequence sampling over contiguous or fragmented trajectories, offline-to-online transitions, generation-stamped slots, operational statistics, and consume-after-sample support. See #3947, #4045, #4051, #4031, and #4180 by @theap06 and @vmoens, #4050 by @coder-jayp, #4046 by @harryfrzz, and #3911 by @vmoens.
  • On-policy and LLM training add GRPOTrainer, PPO-EWMA with a decoupled proximal policy, reusable percentile value normalization, and group-wise value-estimator standardization. See #4140 by @coder-jayp and #4126, #4127, #4212, #4213, and #4214 by @vmoens.
  • Simulation support expands with a CraftGround Minecraft wrapper, MuJoCo contact and site APIs, the MicroDuckEnv multitask locomotion environment and training examples, and richer torchrl.render notebooks and MuJoCo WASM tooling. See #4107, #4183, #4203, #4204, #4205, #4212, and #4218 by @vmoens.
  • The VLA stack provides TensorDict-native VLA data and policy abstractions, LeRobot and OpenVLA integration, and LIBERO training recipes. See #3884, #3995, and the VLA data and LIBERO commit stacks by @vmoens.

MicroDuck in motion

A single multitask policy handles forward walking, backward walking, sidestepping, and hopping. This checkpoint was filmed after 27.95 million transitions with one native simulator per tile; the top row walks forward and backward, while the bottom row sidesteps and hops.

MicroDuck multitask policy: forward, backward, sidestep, and hop

Watch or download the MP4 (0.8 MB). Training details and learning curves are in #4212.

Measured performance

The tables below reproduce measurements reported in the linked pull requests. Hardware, tensor shapes, dependency versions, and benchmark scope differ between tables; comparisons are meaningful within each table, not across them.

Block-GRU recurrence

On an NVIDIA GB200 with BF16 activations, FP32 parameters, batch 16, hidden and projection width 512, eight blocks, and one layer, the fused Triton block-GRU substantially reduces recurrent execution time relative to the scan backend. #4160

Sequence length Backend Forward Forward + backward Peak memory
64 scan 65.426 ms 115.099 ms 0.160 GB
64 Triton 4.223 ms 7.443 ms 0.149 GB
512 scan 540.584 ms 855.082 ms 0.782 GB
512 Triton 28.273 ms 62.608 ms 0.690 GB

A recompute-based Triton backward then reduced saved-state materialization and peak memory on the same hardware; forward timings remained unchanged within noise. #4174

Shape Earlier forward + backward Recompute backward Peak memory
B=256, T=512, H=128, P=64 26.943 ms 19.589 ms 1.612 -> 0.807 GB
B=2048, T=64, H=128, P=64 14.307 ms 6.883 ms 1.614 -> 0.809 GB
B=2048, T=512, H=128, P=64 103.401 ms 43.963 ms 12.453 -> 5.578 GB
B=16, T=512, H=P=512 63.405 ms 62.694 ms 0.690 -> 0.355 GB

RSSM rollout and complete DreamerV3 updates

For batch 16, sequence length 64, and belief width 512 on an NVIDIA A10G, compiled higher-order scan improves RSSM inference and training throughput. Unrolling the scan by eight retains that speed while sharply reducing its training allocation. #4132

RSSM execution Inference Forward + backward Training peak allocation
eager loop 90.1 ms 216.2 ms 43.24 MiB
compiled loop step 49.9 ms 134.6 ms 36.35 MiB
compiled scan, unroll 1 22.78 ms 101.00 ms 210.18 MiB
compiled scan, unroll 8 18.01 ms 93.04 ms 58.89 MiB

At the complete-learner level on an NVIDIA GB200, CUDA-graph capture removes thousands of per-update launches and turns the full model, actor, value, replay-value, optimizer, and slow-target update into the fastest measured path. Compilation and capture warm-up, replay sampling, and environment collection are excluded. #4216

Learner implementation Median per update Transitions/s Relative result
pinned upstream DreamerV3 JAX 22.04 ms 46,468 1.00x
TorchRL compiled scan 358.51 ms 2,856 16.3x slower
TorchRL compiled scan + CUDA graph 17.83 ms 57,415 1.24x faster

Async environments and inference

On an NVIDIA GB200 with 16 environments, 5 ms environment latency, and a two-layer 1024-unit policy, consolidated direct process slots nearly double collection throughput relative to the shared-memory coordinator. Grouping four environments per worker instead targets host-memory efficiency. #4272

Transport Frames/s Median batch latency Host RSS
shared-memory coordinator, one env/worker 1,009.81-1,061.53 240.45-242.73 ms 15,312-15,375 MiB
shared-memory coordinator, four envs/worker 1,072.05-1,077.66 237.55-238.14 ms 5,459-5,969 MiB
direct process slots, consolidated transfer 1,940.39-2,024.66 115.61-122.17 ms 14,763-15,781 MiB

For a 170.8-million-parameter policy on the same GPU class, static CUDA-graph batches raise inference-server request throughput across batch sizes. These are server-only measurements: the integrated A10G collection workload in the same PR remained about 9-10% slower than eager end to end, which identifies collation and transport as separate optimization targets. #4269

Batch Eager requests/s CUDA-graph requests/s Eager median round trip Graph median round trip
1 171.58-183.98 538.48-577.29 5.36-5.75 ms 1.72-1.79 ms
16 1,248.55-1,303.08 3,465.92-3,587.45 11.44-12.03 ms 4.32-4.49 ms
64 1,647.64-1,692.65 5,291.55-5,431.54 32.35-34.10 ms 11.25-11.58 ms

End-to-end asynchronous training

The combined process-slot, pinned-staging, replay write-back, and setup-heap stack was measured on the same 64-environment pixel workload and 4-GPU node for both runs. The train ratio couples environment and learner rates, so both improve together. This is a stack-level result and is not attributed to any one PR. #4305, #4306, #4307, and #4308

Configuration Environment steps/s Learner updates/s
before the stack 361 0.70
with the stack 1,249 2.44
measured speedup 3.46x 3.46x

Finally, paired 9.1-hour runs show the effect of sending at least 64 consecutive transitions per process-worker message. The gap remained between 17% and 20% in every half-hour window. #4315

Process-slot chunking Whole-run steps/s Whole-run updates/s Steady-window steps/s Steps in 9.1 hours
16 transitions/message 1,018 1.94 1,000-1,033 33.3M
64 transitions/message 1,219 2.33 1,198-1,245 39.9M

Breaking changes

  • QValueModule and QValueActor now default to strict action-shape validation. A mismatch with the action spec raises RuntimeError; pass strict_shape="auto" to reshape compatible outputs or strict_shape=False to disable validation. Explicit legacy None follows the new strict behavior. #4179 by @vmoens.
  • The deprecated MultiCollector.postprocs alias has been removed. Use MultiCollector.postproc. #4108 by @vmoens.
  • TorchRL now requires tensordict>=0.14.2,<0.15.0; upgrade TorchRL and TensorDict together. #4108 and #4212 by @vmoens; the updated dependency floor is tracked in #4319.

Deprecations and future defaults

  • The legacy use_ray_service switch is deprecated in favor of explicit service_backend selection and is scheduled for removal in v0.16. ca3261f52 by @vmoens.
  • The default Trainer checkpoint format will move from the legacy backend to torchrl.checkpoint in v0.15, and the compatibility checkpoint argument is scheduled for removal in v0.16. #3967 by @vmoens.
  • DreamerV3's legacy reward-key fallback is deprecated; write logits to the configured reward-logits key before its removal in v0.16. #4066 by @vmoens.
  • Constructing a multiprocessing AsyncEnvPool without an explicit exchange now warns that the default will change from "queue" to "auto" in v0.15. Select "queue", "shm", or "auto" explicitly to pin the desired behavior. #4200 by @vmoens.

Features

Model-based reinforcement learning

  • WorldModel, WorldModelWrapper, and WorldModelLoss provide a general TensorDict-native abstraction for model-based environments, rollouts, and joint objectives. #3783 by @theap06.
  • DreamerV3 now implements corrected symexp two-hot encoding, reward/value decoding, normalized returns in both gradient modes, slow-critic regularization, categorical KL semantics, and continuation training. #4065, #4066, #4067, #4068, #4073, #4074, and #4125 by @vmoens.
  • DreamerV3BlockGRU and DreamerV3BlockGRUCell expose sequence-first block-GRU modules with a Python implementation and an optional Triton backend whose backward pass avoids materializing block-expanded activations. #4158, #4160, and #4174 by @vmoens.
  • DreamerV3MLPConfig exposes DreamerV3's specialized MLP through the trainer Hydra configuration system. #4187 by @bsprenger.
  • A reproducible DreamerV3 DMC Walker implementation, benchmark, pinned JAX comparison, and one-command reproduction script make the reference training path directly runnable and measurable. #4075 and #4133 by @vmoens and #4196 by @theap06.

Distributed training, inference, and operations

  • Configurable service transports and shared-memory inference support reusable communication backends across collectors, inference servers, replay services, and weight updates. #3955 and #3958 by @vmoens.
  • PolicyClientModule, ProcessInferenceServer, structured inference-server configs, behavior-policy version tracking, and max-inflight controls form a complete remote-policy client path. 307eb5a64, 938d0abd1, cd1406e38, and 539aa07d8 by @vmoens.
  • Trainer execution backends add DDP optimization, Ray data-parallel execution, distributed checkpointing, remote off-policy learners, and WeightSyncScheme publication. 311c1b13f, 6e60c472d, 12bef231d, and dc6dcb8a5 by @vmoens.
  • Collectors can be built through a unified backend interface, while AsyncEnvPool gains shared-memory exchange and automatic parallel environments expose worker metadata. #4022 and #4093 by @vmoens and #4109 by @theap06.
  • AsyncEnvPool(exchange="auto") selects shared-memory transport for compatible multiprocessing environments and falls back to queues for dynamic, heterogeneous, or non-tensor schemas; resolved_exchange reports the selected transport. #4200 by @vmoens.
  • Pull-based monitoring now includes ReplayBuffer.stats(), collector statistics across local, multiprocessing, and Ray collectors, plus LoggerMonitor and Every. #4031, #4032, and #4033 by @theap06.

Replay buffers and checkpointing

  • The trajectory query API introduces traj, Trajectory, predicates, filters, and iterators for selecting replay data through composable trajectory expressions. #3947 by @theap06.
  • SampleUnit, Transition, and Sequence separate sampling semantics from storage. Sequence sampling supports exact boundary policies, burn-in, bootstrap context, dilation, per-anchor priorities, and multidimensional storage. #4045, #4051, and #4085 by @theap06 and #4050 by @coder-jayp.
  • SliceSampler(fragmented=True, step_key=...) reconstructs fixed-length logical trajectory slices when trajectories are interleaved or non-contiguous in one-dimensional TensorDict-backed storage. #4180 by @vmoens.
  • OfflineToOnlineReplayBuffer, prefill_replay_buffer, and OfflineToOnlineTrainer support staged transitions from static datasets to online collection, with a runnable SOTA implementation. #3900 and #3904 by @theap06.
  • Round-robin writers can stamp generations to reject stale priority updates, and consuming samplers can retire data after sampling. #4046 by @harryfrzz and #3911 by @vmoens.
  • The new torchrl.checkpoint package handles structured state, adapters, strictness policies, RNG state, best-checkpoint selection, rotation, retention, and dependency-version diagnostics. #3967 by @vmoens and #4064 and #4098 by @theap06.

VLA, simulation, video, and rendering

  • TensorDict-native VLA schemas, robot dataset metadata, action tokenizers, ActionChunkTransform, VLAWrapperBase, TinyVLA, and LeRobotPolicyWrapper cover data preparation through policy execution. b6048dc43, e9a26b5be, a777bc4bc, and #3884 by @vmoens.
  • LeRobot datasets, OpenVLA preprocessing, LIBERO environments, behavior cloning, token PPO, and a VLA-GRPO LIBERO recipe provide an end-to-end robotics training path. 07ac42692, #3995, ed66a0f48, and 775c52748 by @vmoens.
  • CraftGroundEnv and CraftGroundWrapper integrate CraftGround's Minecraft environments with explicit pixel and action specs and composable TorchRL rewards and termination. #4107 by @vmoens.
  • New wrappers integrate MuJoCo Playground and mjlab environments, including a batched PPO example for accelerator-backed simulation. #3751 by @itwasabhi and #3910 by @vmoens.
  • MujocoEnv.geom_contacts() and site_positions() expose backend-independent simulation signals, and patch_xml=False preserves local model paths so native MuJoCo, MJX, and MuJoCo-Torch can resolve relative assets. #4190 and #4204 by @vmoens.
  • MicroDuckEnv provides backend-independent biped locomotion with tensorclass task libraries, weighted reward registration, task sampling, standing, tracking, sidestepping, and jumping presets. Recurrent PPO, MJLab PPO, closed-form gait, checkpoint-backed rendering, and evaluation-video examples cover the training workflow. #4183, #4204, #4205, #4212, and #4218 by @vmoens.
  • The macro-action Cartesian IK API supports partial pose constraints, letting callers control only selected position and orientation components. #3935 by @vmoens.
  • VideoClipRef adds lazy, compact, device-aware video decoding, while torchrl.render can produce videos, checkpoint-backed artifacts, live-parameter notebooks, and MuJoCo WASM viewers with an optional bundled Node runtime. #3824, #3826, #3839, a38ea665c, 54f6d9ccd, and #4203 by @vmoens.

Objectives and LLM workflows

  • RNDLoss, the RND transform, and a PPO+RND MuJoCo implementation add intrinsic-reward training with a complete runnable example. #3889 and #3905 by @theap06.
  • TQCLoss adds truncated quantile critics, list-valued critic updates, and numerical-contract coverage. #4077, #4078, and #4079 by @gtnv.
  • DistillationLoss supports token-level LLM knowledge distillation, and TRL interoperability adapters translate policies, rollouts, and training data between TRL and TorchRL. #4038 by @theap06 and #4070 by @coder-jayp.
  • GRPOTrainer, GRPOOptimizationStepper, and WeightSyncHook provide a modular GRPO training loop with gradient accumulation, mixed precision, clipping, and asynchronous synchronization to distributed LLM inference engines; synchronous and asynchronous SOTA pipelines exercise the complete path. #4140 by @coder-jayp.
  • PercentileValueNorm tracks an exponential moving average of configurable quantiles, while ValueNorm.scale() exposes scale-only normalization. PPO and A2C losses accept advantage_norm for robust, opt-in advantage scaling. #4126 and #4127 by @vmoens.
  • PPO losses support decoupled proximal policies through delay_actor and capped behavior-policy ratios through max_importance_ratio. Trainer and Hydra integration can update the proximal policy with SoftUpdate or HardUpdate, including a runnable PPO-EWMA configuration. #4213 and #4214 by @vmoens.
  • GAE, VTrace, TD0Estimator, TD1Estimator, and TDLambdaEstimator accept group_key to standardize advantages or rewards independently within groups such as tasks or agents. #4212 by @vmoens.
  • Reusable value transforms and configurable chunk dimensions extend value normalization and estimator execution without changing existing call patterns. #4096 by @vmoens and #4003 by @lin-erica.

Asynchronous collection, inference, and replay

  • AsyncBatchedCollector supports compile-safe pausing, background replay collection, direct replay routing, CPU affinity, multi-environment workers, process-hosted inference, and fast asynchronous defaults. #4263, #4266, #4270, #4271, #4272, #4275, and #4313 by @vmoens.
  • Static CUDA-graph inference batches, bounded process-slot collection, chunked worker results, and pinned staging batches reduce synchronization and transfer overhead while preserving request metadata. #4269, #4297, #4303, #4305, and #4306 by @vmoens.
  • Replay operations can be ordered across concurrent producers, and native stream replay supports trajectory-aware sampling and direct asynchronous collection. #4273 and #4274 by @vmoens.
  • Hydra configurations expose replay-buffer transport and multi-collector construction options. #4254 and #4291 by @aswanth-07.

DreamerV3 pipeline composition

  • Image encoders and decoders, discrete-action RSSM support, actor entropy reporting, the optimizer, composed reconstruction heads, model loss, optimization stepper, and acting-state estimation are available as reusable public components. #4277, #4278, #4282, #4283, #4284, and #4285 by @vmoens.
  • Environment and execution controls, native replay checkpoints, explicit exploration, stored episode flags, and cross-episode replay sequences complete the configurable asynchronous training path. #4280, #4281, and #4310 by @vmoens.
  • DreamerV3 two-hot decoding is reduction-safe under compiler contraction, and RSSM scan carries keep stable dtypes under autocast. #4290 by @theap06.

Environments and integrations

  • LBForagingEnv integrates Level-Based Foraging through the standard TorchRL environment interface. #4301 by @Iliamsou.

Public API additions

The feature sections above describe the complete workflows. This index collects the principal public entry points for readers upgrading code or exploring the API:

  • Model-based RL: WorldModel, WorldModelWrapper, WorldModelLoss, DreamerV3BlockGRU, DreamerV3BlockGRUCell, and DreamerV3MLPConfig. #3783, #4158, #4160, #4174, and #4187.
  • Replay and checkpointing: traj, Trajectory, SampleUnit, Transition, Sequence, OfflineToOnlineReplayBuffer, prefill_replay_buffer, and the torchrl.checkpoint package. #3947, #4045, #3900, and #3967.
  • Distributed execution and observability: PolicyClientModule, ProcessInferenceServer, WeightSyncScheme, ReplayBuffer.stats(), LoggerMonitor, and Every.
  • Collection and inference: AsyncBatchedCollector, process-hosted inference, background replay collection, native stream replay, and AsyncEnvPool.resolved_exchange. #4263, #4270, #4274, and #4200.
  • Objectives and trainers: RNDLoss, TQCLoss, DistillationLoss, GRPOTrainer, GRPOOptimizationStepper, WeightSyncHook, and PercentileValueNorm. #3889, #4077, #4038, #4140, and #4126.
  • Simulation and environments: CraftGroundEnv, CraftGroundWrapper, MicroDuckEnv, LBForagingEnv, MujocoEnv.geom_contacts(), and MujocoEnv.site_positions(). #4107, #4183, #4301, and #4190.
  • VLA, data, and rendering: VideoClipRef, ActionChunkTransform, VLAWrapperBase, TinyVLA, and LeRobotPolicyWrapper, together with the torchrl.render artifact and notebook workflows. #3824 and #3884.

Performance

  • Recurrent kernels gain tiled large-hidden support, 64-bit offsets, recompute-based backward passes, narrower canonicalization, specialized GRU backward scans, and hardened Triton execution. #3818, #3752, #4154, and #4175 by @vmoens.
  • DreamerV3 RSSM rollout avoids unnecessary Python and tensor work, and the block-GRU scan adds optimized recurrence plus a Triton backward pass that reduces runtime and peak activation memory. #4132, #4159, #4160, and #4174 by @vmoens.
  • The DreamerV3 DMC reproduction supports an opt-in CUDA-graph learner path for fixed-shape updates. On the reference GB200 benchmark, complete learner updates improve from 358.51 ms with the compiled scan to 17.83 ms with CUDA graphs, excluding compilation and capture warm-up. #4216 by @vmoens.
  • MuJoCo-Torch batched physics steps compile as one full graph by default when compile_step=True, preventing silent eager fallback after graph breaks or recompilation limits. #4202 by @vmoens.
  • MultiSyncCollector preemption now uses a shared-memory interruptor instead of busy-waiting, with macOS support, and stops gathering as soon as all worker outputs arrive. #3843 and #3841 by @vmoens.
  • Replay-boundary queries are cached and vectorized, TransformedEnv specs lock by default, and parent-side LIBERO construction is eliminated. #4084 by @theap06, #4121 by @mathieuorhan, and #4028 by @vmoens.

  • Async environment and collector execution caches setup metadata, bounds shared mappings, batches shared-memory coordination, and reduces process-slot transfer overhead. #4258, #4261, #4264, #4305, and #4306 by @vmoens.

  • DreamerV3 uses native replay collection, compiles the complete learner step, selects its compile strategy automatically, keeps replay write-back off the learner stream, and freezes the setup heap before training. #4265, #4268, #4302, #4307, and #4308 by @vmoens.
  • Process-slot collectors send at least 64 consecutive transitions per worker message by default, improving driver throughput while preserving an explicit low-latency override. #4315 by @vmoens.

Bug fixes

Objectives, estimators, and models

  • Objective and distribution fixes cover saturated TanhNormal scoring, multidimensional DiffusionActor training, sparse masked actions, singleton IQL gradients, PPO execution, lazy target initialization, TruncatedNormal dtype preservation, DreamerV3 reparameterized return normalization, shared MLP output bias, V-trace bootstrapping at in-batch truncations, compile-friendly analytic entropy and KL selection, and device-safe percentile quantiles. #4080, #4148, #4149, and #4150 by @gtnv, #4114 by @theap06, #4113 and #4137 by @aswanth-07, #4125 by @vmoens, #4169 by @quinnarnold, #4210 by @yurekami, #4234 by @YeonwooSung, and #4241 by @coder-jayp.

Collectors, environments, and specs

  • Collector and environment fixes preserve renamed done keys, caller CUDA streams, shared-device weight synchronization, dynamic specs, and MPS-safe ParallelEnv transport. Compiled stochastic policies keep fresh random samples under CUDA graphs, async receive deadlines bound the complete call without losing partial results, queue-backed pools shut down without deadlocking, stacked composite specs propagate lock state, and batched native MuJoCo workers receive distinct reproducible seeds. #4147 and #4146 by @gtnv, #4027 by @theap06, #4143 by @coder-jayp, #3867 by @discobot, and #3873, #4184, #4197, #4198, and #4208 by @vmoens.

Replay buffers

  • Replay-buffer fixes correct prefetch depth and checkpointing, priority transforms and annealing, sampler restoration, generation validation, root/nested device metadata, and direct slice-indexed storage writes. #3871 by @Agade09, #3925, #3990, #4043, and #4089 by @theap06, and #3914, #3915, #4014, and #4195 by @vmoens.

Training and checkpointing

  • Trainer checkpoints now persist optimizer state across direct optimizers, default optimization steppers, and optimizer hooks, and trainer imports select the compatible GradScaler location on PyTorch releases older than 2.3. #4189 by @bsprenger and #4244 by @coder-jayp.

LLM, datasets, and optional integrations

  • LLM and dataset fixes normalize reward losses over valid pairs, infer tokenizer padding correctly, preserve real attention masks for padded single-string tokenization, follow module device moves in SFT and distillation losses, reject empty distillation sequences, keep the vLLM FP32 plugin opt-in, and scope Minari dataset paths correctly. #3886 and #3887 by @fallintoplace, #4082 by @aswanth-07, #3868 by @vmoens, #3920 by @younik, and #4237 and #4238 by @YeonwooSung.

Configuration and environment APIs

  • Hydra configuration targets for KLRewardTransformConfig and StorageEnsembleWriterConfig now resolve to importable implementations with matching public fields. #4236 by @YeonwooSung.
  • ChessEnv.rand_action(tensordict) now samples from the action mask represented by the supplied mask or FEN/PGN state, including through the public transformed environment. #4239 by @YeonwooSung.

Logging and asynchronous execution

  • WandbLogger forwards its configured base URL during authentication, and runtime diagnostics report CUDA timing and synchronization context more precisely. #4259 and #4260 by @vmoens.
  • Async environment construction remains compile-safe for non-tensor data, while process-slot collection is bounded and reliably stops orphaned workers. #4276 and #4297 by @vmoens.
  • MultiCategorical.project() clamps negative inputs before dtype conversion, and TensorDictPrimer preserves integer and boolean spec dtypes during reset and step. #4317 by @yurekami and #4318 by @theap06.

Documentation and validation

Tutorials and guides

  • Tutorials now use current replay-buffer storage, CUDA-graph, multi-agent reset, transform, and TRL interoperability patterns; their multiprocessing setup also works on Windows. The recurrent DQN material clarifies the required TensorDictPrimer state lifecycle, and the MicroDuck tutorial covers task libraries, reward registration, task sampling, per-task advantages, backends, gait control, and training. #4181 and #4219 by @vmoens, #4123 by @coder-jayp, and #4220 by @YeonwooSung.

API reference

  • The value-objective reference documents PQN lambda returns through DQNLoss with ValueEstimators.TDLambda, including greedy next-action bootstrapping. #4233 by @YeonwooSung.

Examples, tests, benchmarks, and CI

  • Examples have systematic CI execution, AsyncEnvPool has continuous dispatch and throughput benchmarks, and regression coverage was strengthened for noisy layers, SFT loss, reward-to-go values, tensor lambdas, chat-template round trips, optional RSSM scan dependencies, and repository-independent render test paths. #4152 and #4182 by @vmoens, #4163, #4168, #4170, and #4171 by @quinnarnold, #4235 by @YeonwooSung, and #4242 and #4243 by @coder-jayp.
  • Release maintenance refreshes the expert-iteration NLTK and Transformers dependencies, records the repository contribution contracts, and exercises MuJoCo-Torch from main alongside nightly PyTorch on macOS. #4199, #4201, #4207, and #4211.

  • Continuous benchmarks track async collection across merges, direct process-slot collection, benchmarked commit identities, and baseline collector and DreamerV3 measurements. Performance alerts no longer prevent result publication, and labelled merges can run the complete benchmark suite. #4287, #4289, #4292, #4293, #4298, and #4312 by @vmoens.

  • The DreamerV3 process-inference benchmark selects the transport-owned environment exchange so both thread and process series remain runnable. #4320 by @vmoens.
  • Regression coverage initializes MicroDuck fixtures deterministically, exercises prioritized-replay weights in DDPG, and preserves metadata in static inference batches. #4279 by @coder-jayp and #4288 and #4303 by @vmoens.
  • Contributor guidance records target-device tensor and module construction as a repository contract. #4314 by @vmoens.

Upgrade notes

  • Install matching release lines together: pip install "torchrl==0.14.0" "tensordict>=0.14.2,<0.15.0".
  • Review code using QValueModule or QValueActor with shaped categorical specs. Choose strict validation, automatic reshaping, or explicit opt-out rather than relying on the former warning-only default.
  • Replace MultiCollector.postprocs with MultiCollector.postproc, and review the v0.15 checkpoint and collector-default warnings before the next release.
  • Pass an explicit exchange to multiprocessing AsyncEnvPool instances to keep queue transport or opt into automatic shared-memory selection before the default changes.

Full changelog

For the complete commit-by-commit history, see v0.13.3...v0.14.0.

Project stewardship

We are delighted to announce that @theap06 is joining @vmoens as a maintainer of TorchRL. His exceptional work across model-based reinforcement learning, replay buffers, checkpointing, and the library as a whole has made TorchRL substantially stronger. Thank you, @theap06, for the care, energy, and extraordinary amount of work you have brought to the project. We are looking forward to building and shipping what comes next together, alongside the entire TorchRL community.

Contributors

Thanks to @Agade09, @aswanth-07, @bsprenger, @coder-jayp, @discobot, @fallintoplace, @gtnv, @harryfrzz, @Iliamsou, @itwasabhi, @lin-erica, @mathieuorhan, @ParamThakkar123, @quinnarnold, @saputkin, @theap06, @unography, @vmoens, @xyf5432, @YeonwooSung, @younik, and @yurekami for the changes included in this release.

We especially welcome first-time contributors @Agade09, @aswanth-07, @discobot, @fallintoplace, @gtnv, @harryfrzz, @Iliamsou, @quinnarnold, @saputkin, @unography, @xyf5432, @YeonwooSung, and @yurekami.

Source: README.md, updated 2026-09-10