TorchRL v0.14.0
TorchRL 0.14 is a major feature release spanning model-based reinforcement learning, distributed training and inference, replay-buffer composition, checkpointing, on-policy optimization, vision-language-action workflows, and simulation tooling. It also tightens several contracts announced in earlier releases and requires TensorDict 0.14.2 or later in the 0.14 release line.
Highlights
- DreamerV3 receives an end-to-end correctness and performance overhaul, including a block-RSSM core, public
DreamerV3BlockGRUmodules, Triton and CUDA-graph acceleration, corrected losses and continuation semantics, reusable percentile return normalization, Hydra-compatible model configuration, and a reproducible DMC Walker benchmark. See #4072, #4074, #4075, #4125, #4126, #4132, #4133, #4158, #4160, #4174, #4187, #4196, and #4216 by @vmoens, @theap06, and @bsprenger, with correctness fixes in #4119 by @gtnv and #4118 and #4128 by @unography. - Distributed execution now covers configurable TorchRL service transports, shared-memory inference, process inference servers, DDP and Ray trainer execution, remote learners, and unified collector backend construction. See #3955, #3958, #4022, and the distributed trainer commit stack by @vmoens.
- Replay buffers gain a trajectory query language, composable sample units, rich sequence sampling over contiguous or fragmented trajectories, offline-to-online transitions, generation-stamped slots, operational statistics, and consume-after-sample support. See #3947, #4045, #4051, #4031, and #4180 by @theap06 and @vmoens, #4050 by @coder-jayp, #4046 by @harryfrzz, and #3911 by @vmoens.
- On-policy and LLM training add
GRPOTrainer, PPO-EWMA with a decoupled proximal policy, reusable percentile value normalization, and group-wise value-estimator standardization. See #4140 by @coder-jayp and #4126, #4127, #4212, #4213, and #4214 by @vmoens. - Simulation support expands with a CraftGround Minecraft wrapper, MuJoCo contact and site APIs, the
MicroDuckEnvmultitask locomotion environment and training examples, and richertorchrl.rendernotebooks and MuJoCo WASM tooling. See #4107, #4183, #4203, #4204, #4205, #4212, and #4218 by @vmoens. - The VLA stack provides TensorDict-native VLA data and policy abstractions, LeRobot and OpenVLA integration, and LIBERO training recipes. See #3884, #3995, and the VLA data and LIBERO commit stacks by @vmoens.
MicroDuck in motion
A single multitask policy handles forward walking, backward walking, sidestepping, and hopping. This checkpoint was filmed after 27.95 million transitions with one native simulator per tile; the top row walks forward and backward, while the bottom row sidesteps and hops.

Watch or download the MP4 (0.8 MB). Training details and learning curves are in #4212.
Measured performance
The tables below reproduce measurements reported in the linked pull requests. Hardware, tensor shapes, dependency versions, and benchmark scope differ between tables; comparisons are meaningful within each table, not across them.
Block-GRU recurrence
On an NVIDIA GB200 with BF16 activations, FP32 parameters, batch 16, hidden and projection width 512, eight blocks, and one layer, the fused Triton block-GRU substantially reduces recurrent execution time relative to the scan backend. #4160
| Sequence length | Backend | Forward | Forward + backward | Peak memory |
|---|---|---|---|---|
| 64 | scan | 65.426 ms | 115.099 ms | 0.160 GB |
| 64 | Triton | 4.223 ms | 7.443 ms | 0.149 GB |
| 512 | scan | 540.584 ms | 855.082 ms | 0.782 GB |
| 512 | Triton | 28.273 ms | 62.608 ms | 0.690 GB |
A recompute-based Triton backward then reduced saved-state materialization and peak memory on the same hardware; forward timings remained unchanged within noise. #4174
| Shape | Earlier forward + backward | Recompute backward | Peak memory |
|---|---|---|---|
| B=256, T=512, H=128, P=64 | 26.943 ms | 19.589 ms | 1.612 -> 0.807 GB |
| B=2048, T=64, H=128, P=64 | 14.307 ms | 6.883 ms | 1.614 -> 0.809 GB |
| B=2048, T=512, H=128, P=64 | 103.401 ms | 43.963 ms | 12.453 -> 5.578 GB |
| B=16, T=512, H=P=512 | 63.405 ms | 62.694 ms | 0.690 -> 0.355 GB |
RSSM rollout and complete DreamerV3 updates
For batch 16, sequence length 64, and belief width 512 on an NVIDIA A10G, compiled higher-order scan improves RSSM inference and training throughput. Unrolling the scan by eight retains that speed while sharply reducing its training allocation. #4132
| RSSM execution | Inference | Forward + backward | Training peak allocation |
|---|---|---|---|
| eager loop | 90.1 ms | 216.2 ms | 43.24 MiB |
| compiled loop step | 49.9 ms | 134.6 ms | 36.35 MiB |
| compiled scan, unroll 1 | 22.78 ms | 101.00 ms | 210.18 MiB |
| compiled scan, unroll 8 | 18.01 ms | 93.04 ms | 58.89 MiB |
At the complete-learner level on an NVIDIA GB200, CUDA-graph capture removes thousands of per-update launches and turns the full model, actor, value, replay-value, optimizer, and slow-target update into the fastest measured path. Compilation and capture warm-up, replay sampling, and environment collection are excluded. #4216
| Learner implementation | Median per update | Transitions/s | Relative result |
|---|---|---|---|
| pinned upstream DreamerV3 JAX | 22.04 ms | 46,468 | 1.00x |
| TorchRL compiled scan | 358.51 ms | 2,856 | 16.3x slower |
| TorchRL compiled scan + CUDA graph | 17.83 ms | 57,415 | 1.24x faster |
Async environments and inference
On an NVIDIA GB200 with 16 environments, 5 ms environment latency, and a two-layer 1024-unit policy, consolidated direct process slots nearly double collection throughput relative to the shared-memory coordinator. Grouping four environments per worker instead targets host-memory efficiency. #4272
| Transport | Frames/s | Median batch latency | Host RSS |
|---|---|---|---|
| shared-memory coordinator, one env/worker | 1,009.81-1,061.53 | 240.45-242.73 ms | 15,312-15,375 MiB |
| shared-memory coordinator, four envs/worker | 1,072.05-1,077.66 | 237.55-238.14 ms | 5,459-5,969 MiB |
| direct process slots, consolidated transfer | 1,940.39-2,024.66 | 115.61-122.17 ms | 14,763-15,781 MiB |
For a 170.8-million-parameter policy on the same GPU class, static CUDA-graph batches raise inference-server request throughput across batch sizes. These are server-only measurements: the integrated A10G collection workload in the same PR remained about 9-10% slower than eager end to end, which identifies collation and transport as separate optimization targets. #4269
| Batch | Eager requests/s | CUDA-graph requests/s | Eager median round trip | Graph median round trip |
|---|---|---|---|---|
| 1 | 171.58-183.98 | 538.48-577.29 | 5.36-5.75 ms | 1.72-1.79 ms |
| 16 | 1,248.55-1,303.08 | 3,465.92-3,587.45 | 11.44-12.03 ms | 4.32-4.49 ms |
| 64 | 1,647.64-1,692.65 | 5,291.55-5,431.54 | 32.35-34.10 ms | 11.25-11.58 ms |
End-to-end asynchronous training
The combined process-slot, pinned-staging, replay write-back, and setup-heap stack was measured on the same 64-environment pixel workload and 4-GPU node for both runs. The train ratio couples environment and learner rates, so both improve together. This is a stack-level result and is not attributed to any one PR. #4305, #4306, #4307, and #4308
| Configuration | Environment steps/s | Learner updates/s |
|---|---|---|
| before the stack | 361 | 0.70 |
| with the stack | 1,249 | 2.44 |
| measured speedup | 3.46x | 3.46x |
Finally, paired 9.1-hour runs show the effect of sending at least 64 consecutive transitions per process-worker message. The gap remained between 17% and 20% in every half-hour window. #4315
| Process-slot chunking | Whole-run steps/s | Whole-run updates/s | Steady-window steps/s | Steps in 9.1 hours |
|---|---|---|---|---|
| 16 transitions/message | 1,018 | 1.94 | 1,000-1,033 | 33.3M |
| 64 transitions/message | 1,219 | 2.33 | 1,198-1,245 | 39.9M |
Breaking changes
QValueModuleandQValueActornow default to strict action-shape validation. A mismatch with the action spec raisesRuntimeError; passstrict_shape="auto"to reshape compatible outputs orstrict_shape=Falseto disable validation. Explicit legacyNonefollows the new strict behavior. #4179 by @vmoens.- The deprecated
MultiCollector.postprocsalias has been removed. UseMultiCollector.postproc. #4108 by @vmoens. - TorchRL now requires
tensordict>=0.14.2,<0.15.0; upgrade TorchRL and TensorDict together. #4108 and #4212 by @vmoens; the updated dependency floor is tracked in #4319.
Deprecations and future defaults
- The legacy
use_ray_serviceswitch is deprecated in favor of explicitservice_backendselection and is scheduled for removal in v0.16. ca3261f52 by @vmoens. - The default
Trainercheckpoint format will move from the legacy backend totorchrl.checkpointin v0.15, and the compatibility checkpoint argument is scheduled for removal in v0.16. #3967 by @vmoens. - DreamerV3's legacy reward-key fallback is deprecated; write logits to the configured reward-logits key before its removal in v0.16. #4066 by @vmoens.
- Constructing a multiprocessing
AsyncEnvPoolwithout an explicitexchangenow warns that the default will change from"queue"to"auto"in v0.15. Select"queue","shm", or"auto"explicitly to pin the desired behavior. #4200 by @vmoens.
Features
Model-based reinforcement learning
WorldModel,WorldModelWrapper, andWorldModelLossprovide a general TensorDict-native abstraction for model-based environments, rollouts, and joint objectives. #3783 by @theap06.- DreamerV3 now implements corrected symexp two-hot encoding, reward/value decoding, normalized returns in both gradient modes, slow-critic regularization, categorical KL semantics, and continuation training. #4065, #4066, #4067, #4068, #4073, #4074, and #4125 by @vmoens.
DreamerV3BlockGRUandDreamerV3BlockGRUCellexpose sequence-first block-GRU modules with a Python implementation and an optional Triton backend whose backward pass avoids materializing block-expanded activations. #4158, #4160, and #4174 by @vmoens.DreamerV3MLPConfigexposes DreamerV3's specialized MLP through the trainer Hydra configuration system. #4187 by @bsprenger.- A reproducible DreamerV3 DMC Walker implementation, benchmark, pinned JAX comparison, and one-command reproduction script make the reference training path directly runnable and measurable. #4075 and #4133 by @vmoens and #4196 by @theap06.
Distributed training, inference, and operations
- Configurable service transports and shared-memory inference support reusable communication backends across collectors, inference servers, replay services, and weight updates. #3955 and #3958 by @vmoens.
PolicyClientModule,ProcessInferenceServer, structured inference-server configs, behavior-policy version tracking, and max-inflight controls form a complete remote-policy client path. 307eb5a64, 938d0abd1, cd1406e38, and 539aa07d8 by @vmoens.- Trainer execution backends add DDP optimization, Ray data-parallel execution, distributed checkpointing, remote off-policy learners, and
WeightSyncSchemepublication. 311c1b13f, 6e60c472d, 12bef231d, and dc6dcb8a5 by @vmoens. - Collectors can be built through a unified backend interface, while
AsyncEnvPoolgains shared-memory exchange and automatic parallel environments expose worker metadata. #4022 and #4093 by @vmoens and #4109 by @theap06. AsyncEnvPool(exchange="auto")selects shared-memory transport for compatible multiprocessing environments and falls back to queues for dynamic, heterogeneous, or non-tensor schemas;resolved_exchangereports the selected transport. #4200 by @vmoens.- Pull-based monitoring now includes
ReplayBuffer.stats(), collector statistics across local, multiprocessing, and Ray collectors, plusLoggerMonitorandEvery. #4031, #4032, and #4033 by @theap06.
Replay buffers and checkpointing
- The trajectory query API introduces
traj,Trajectory, predicates, filters, and iterators for selecting replay data through composable trajectory expressions. #3947 by @theap06. SampleUnit,Transition, andSequenceseparate sampling semantics from storage. Sequence sampling supports exact boundary policies, burn-in, bootstrap context, dilation, per-anchor priorities, and multidimensional storage. #4045, #4051, and #4085 by @theap06 and #4050 by @coder-jayp.SliceSampler(fragmented=True, step_key=...)reconstructs fixed-length logical trajectory slices when trajectories are interleaved or non-contiguous in one-dimensional TensorDict-backed storage. #4180 by @vmoens.OfflineToOnlineReplayBuffer,prefill_replay_buffer, andOfflineToOnlineTrainersupport staged transitions from static datasets to online collection, with a runnable SOTA implementation. #3900 and #3904 by @theap06.- Round-robin writers can stamp generations to reject stale priority updates, and consuming samplers can retire data after sampling. #4046 by @harryfrzz and #3911 by @vmoens.
- The new
torchrl.checkpointpackage handles structured state, adapters, strictness policies, RNG state, best-checkpoint selection, rotation, retention, and dependency-version diagnostics. #3967 by @vmoens and #4064 and #4098 by @theap06.
VLA, simulation, video, and rendering
- TensorDict-native VLA schemas, robot dataset metadata, action tokenizers,
ActionChunkTransform,VLAWrapperBase,TinyVLA, andLeRobotPolicyWrappercover data preparation through policy execution. b6048dc43, e9a26b5be, a777bc4bc, and #3884 by @vmoens. - LeRobot datasets, OpenVLA preprocessing, LIBERO environments, behavior cloning, token PPO, and a VLA-GRPO LIBERO recipe provide an end-to-end robotics training path. 07ac42692, #3995, ed66a0f48, and 775c52748 by @vmoens.
CraftGroundEnvandCraftGroundWrapperintegrate CraftGround's Minecraft environments with explicit pixel and action specs and composable TorchRL rewards and termination. #4107 by @vmoens.- New wrappers integrate MuJoCo Playground and mjlab environments, including a batched PPO example for accelerator-backed simulation. #3751 by @itwasabhi and #3910 by @vmoens.
MujocoEnv.geom_contacts()andsite_positions()expose backend-independent simulation signals, andpatch_xml=Falsepreserves local model paths so native MuJoCo, MJX, and MuJoCo-Torch can resolve relative assets. #4190 and #4204 by @vmoens.MicroDuckEnvprovides backend-independent biped locomotion with tensorclass task libraries, weighted reward registration, task sampling, standing, tracking, sidestepping, and jumping presets. Recurrent PPO, MJLab PPO, closed-form gait, checkpoint-backed rendering, and evaluation-video examples cover the training workflow. #4183, #4204, #4205, #4212, and #4218 by @vmoens.- The macro-action Cartesian IK API supports partial pose constraints, letting callers control only selected position and orientation components. #3935 by @vmoens.
VideoClipRefadds lazy, compact, device-aware video decoding, whiletorchrl.rendercan produce videos, checkpoint-backed artifacts, live-parameter notebooks, and MuJoCo WASM viewers with an optional bundled Node runtime. #3824, #3826, #3839, a38ea665c, 54f6d9ccd, and #4203 by @vmoens.
Objectives and LLM workflows
RNDLoss, the RND transform, and a PPO+RND MuJoCo implementation add intrinsic-reward training with a complete runnable example. #3889 and #3905 by @theap06.TQCLossadds truncated quantile critics, list-valued critic updates, and numerical-contract coverage. #4077, #4078, and #4079 by @gtnv.DistillationLosssupports token-level LLM knowledge distillation, and TRL interoperability adapters translate policies, rollouts, and training data between TRL and TorchRL. #4038 by @theap06 and #4070 by @coder-jayp.GRPOTrainer,GRPOOptimizationStepper, andWeightSyncHookprovide a modular GRPO training loop with gradient accumulation, mixed precision, clipping, and asynchronous synchronization to distributed LLM inference engines; synchronous and asynchronous SOTA pipelines exercise the complete path. #4140 by @coder-jayp.PercentileValueNormtracks an exponential moving average of configurable quantiles, whileValueNorm.scale()exposes scale-only normalization. PPO and A2C losses acceptadvantage_normfor robust, opt-in advantage scaling. #4126 and #4127 by @vmoens.- PPO losses support decoupled proximal policies through
delay_actorand capped behavior-policy ratios throughmax_importance_ratio. Trainer and Hydra integration can update the proximal policy withSoftUpdateorHardUpdate, including a runnable PPO-EWMA configuration. #4213 and #4214 by @vmoens. GAE,VTrace,TD0Estimator,TD1Estimator, andTDLambdaEstimatoracceptgroup_keyto standardize advantages or rewards independently within groups such as tasks or agents. #4212 by @vmoens.- Reusable value transforms and configurable chunk dimensions extend value normalization and estimator execution without changing existing call patterns. #4096 by @vmoens and #4003 by @lin-erica.
Asynchronous collection, inference, and replay
AsyncBatchedCollectorsupports compile-safe pausing, background replay collection, direct replay routing, CPU affinity, multi-environment workers, process-hosted inference, and fast asynchronous defaults. #4263, #4266, #4270, #4271, #4272, #4275, and #4313 by @vmoens.- Static CUDA-graph inference batches, bounded process-slot collection, chunked worker results, and pinned staging batches reduce synchronization and transfer overhead while preserving request metadata. #4269, #4297, #4303, #4305, and #4306 by @vmoens.
- Replay operations can be ordered across concurrent producers, and native stream replay supports trajectory-aware sampling and direct asynchronous collection. #4273 and #4274 by @vmoens.
- Hydra configurations expose replay-buffer transport and multi-collector construction options. #4254 and #4291 by @aswanth-07.
DreamerV3 pipeline composition
- Image encoders and decoders, discrete-action RSSM support, actor entropy reporting, the optimizer, composed reconstruction heads, model loss, optimization stepper, and acting-state estimation are available as reusable public components. #4277, #4278, #4282, #4283, #4284, and #4285 by @vmoens.
- Environment and execution controls, native replay checkpoints, explicit exploration, stored episode flags, and cross-episode replay sequences complete the configurable asynchronous training path. #4280, #4281, and #4310 by @vmoens.
- DreamerV3 two-hot decoding is reduction-safe under compiler contraction, and RSSM scan carries keep stable dtypes under autocast. #4290 by @theap06.
Environments and integrations
LBForagingEnvintegrates Level-Based Foraging through the standard TorchRL environment interface. #4301 by @Iliamsou.
Public API additions
The feature sections above describe the complete workflows. This index collects the principal public entry points for readers upgrading code or exploring the API:
- Model-based RL:
WorldModel,WorldModelWrapper,WorldModelLoss,DreamerV3BlockGRU,DreamerV3BlockGRUCell, andDreamerV3MLPConfig. #3783, #4158, #4160, #4174, and #4187. - Replay and checkpointing:
traj,Trajectory,SampleUnit,Transition,Sequence,OfflineToOnlineReplayBuffer,prefill_replay_buffer, and thetorchrl.checkpointpackage. #3947, #4045, #3900, and #3967. - Distributed execution and observability:
PolicyClientModule,ProcessInferenceServer,WeightSyncScheme,ReplayBuffer.stats(),LoggerMonitor, andEvery. - Collection and inference:
AsyncBatchedCollector, process-hosted inference, background replay collection, native stream replay, andAsyncEnvPool.resolved_exchange. #4263, #4270, #4274, and #4200. - Objectives and trainers:
RNDLoss,TQCLoss,DistillationLoss,GRPOTrainer,GRPOOptimizationStepper,WeightSyncHook, andPercentileValueNorm. #3889, #4077, #4038, #4140, and #4126. - Simulation and environments:
CraftGroundEnv,CraftGroundWrapper,MicroDuckEnv,LBForagingEnv,MujocoEnv.geom_contacts(), andMujocoEnv.site_positions(). #4107, #4183, #4301, and #4190. - VLA, data, and rendering:
VideoClipRef,ActionChunkTransform,VLAWrapperBase,TinyVLA, andLeRobotPolicyWrapper, together with thetorchrl.renderartifact and notebook workflows. #3824 and #3884.
Performance
- Recurrent kernels gain tiled large-hidden support, 64-bit offsets, recompute-based backward passes, narrower canonicalization, specialized GRU backward scans, and hardened Triton execution. #3818, #3752, #4154, and #4175 by @vmoens.
- DreamerV3 RSSM rollout avoids unnecessary Python and tensor work, and the block-GRU scan adds optimized recurrence plus a Triton backward pass that reduces runtime and peak activation memory. #4132, #4159, #4160, and #4174 by @vmoens.
- The DreamerV3 DMC reproduction supports an opt-in CUDA-graph learner path for fixed-shape updates. On the reference GB200 benchmark, complete learner updates improve from 358.51 ms with the compiled scan to 17.83 ms with CUDA graphs, excluding compilation and capture warm-up. #4216 by @vmoens.
- MuJoCo-Torch batched physics steps compile as one full graph by default when
compile_step=True, preventing silent eager fallback after graph breaks or recompilation limits. #4202 by @vmoens. MultiSyncCollectorpreemption now uses a shared-memory interruptor instead of busy-waiting, with macOS support, and stops gathering as soon as all worker outputs arrive. #3843 and #3841 by @vmoens.-
Replay-boundary queries are cached and vectorized,
TransformedEnvspecs lock by default, and parent-side LIBERO construction is eliminated. #4084 by @theap06, #4121 by @mathieuorhan, and #4028 by @vmoens. -
Async environment and collector execution caches setup metadata, bounds shared mappings, batches shared-memory coordination, and reduces process-slot transfer overhead. #4258, #4261, #4264, #4305, and #4306 by @vmoens.
- DreamerV3 uses native replay collection, compiles the complete learner step, selects its compile strategy automatically, keeps replay write-back off the learner stream, and freezes the setup heap before training. #4265, #4268, #4302, #4307, and #4308 by @vmoens.
- Process-slot collectors send at least 64 consecutive transitions per worker message by default, improving driver throughput while preserving an explicit low-latency override. #4315 by @vmoens.
Bug fixes
Objectives, estimators, and models
- Objective and distribution fixes cover saturated
TanhNormalscoring, multidimensionalDiffusionActortraining, sparse masked actions, singleton IQL gradients, PPO execution, lazy target initialization,TruncatedNormaldtype preservation, DreamerV3 reparameterized return normalization, shared MLP output bias, V-trace bootstrapping at in-batch truncations, compile-friendly analytic entropy and KL selection, and device-safe percentile quantiles. #4080, #4148, #4149, and #4150 by @gtnv, #4114 by @theap06, #4113 and #4137 by @aswanth-07, #4125 by @vmoens, #4169 by @quinnarnold, #4210 by @yurekami, #4234 by @YeonwooSung, and #4241 by @coder-jayp.
Collectors, environments, and specs
- Collector and environment fixes preserve renamed done keys, caller CUDA streams, shared-device weight synchronization, dynamic specs, and MPS-safe
ParallelEnvtransport. Compiled stochastic policies keep fresh random samples under CUDA graphs, async receive deadlines bound the complete call without losing partial results, queue-backed pools shut down without deadlocking, stacked composite specs propagate lock state, and batched native MuJoCo workers receive distinct reproducible seeds. #4147 and #4146 by @gtnv, #4027 by @theap06, #4143 by @coder-jayp, #3867 by @discobot, and #3873, #4184, #4197, #4198, and #4208 by @vmoens.
Replay buffers
- Replay-buffer fixes correct prefetch depth and checkpointing, priority transforms and annealing, sampler restoration, generation validation, root/nested device metadata, and direct slice-indexed storage writes. #3871 by @Agade09, #3925, #3990, #4043, and #4089 by @theap06, and #3914, #3915, #4014, and #4195 by @vmoens.
Training and checkpointing
- Trainer checkpoints now persist optimizer state across direct optimizers, default optimization steppers, and optimizer hooks, and trainer imports select the compatible
GradScalerlocation on PyTorch releases older than 2.3. #4189 by @bsprenger and #4244 by @coder-jayp.
LLM, datasets, and optional integrations
- LLM and dataset fixes normalize reward losses over valid pairs, infer tokenizer padding correctly, preserve real attention masks for padded single-string tokenization, follow module device moves in SFT and distillation losses, reject empty distillation sequences, keep the vLLM FP32 plugin opt-in, and scope Minari dataset paths correctly. #3886 and #3887 by @fallintoplace, #4082 by @aswanth-07, #3868 by @vmoens, #3920 by @younik, and #4237 and #4238 by @YeonwooSung.
Configuration and environment APIs
- Hydra configuration targets for
KLRewardTransformConfigandStorageEnsembleWriterConfignow resolve to importable implementations with matching public fields. #4236 by @YeonwooSung. ChessEnv.rand_action(tensordict)now samples from the action mask represented by the supplied mask or FEN/PGN state, including through the public transformed environment. #4239 by @YeonwooSung.
Logging and asynchronous execution
WandbLoggerforwards its configured base URL during authentication, and runtime diagnostics report CUDA timing and synchronization context more precisely. #4259 and #4260 by @vmoens.- Async environment construction remains compile-safe for non-tensor data, while process-slot collection is bounded and reliably stops orphaned workers. #4276 and #4297 by @vmoens.
- MultiCategorical.project() clamps negative inputs before dtype conversion, and TensorDictPrimer preserves integer and boolean spec dtypes during reset and step. #4317 by @yurekami and #4318 by @theap06.
Documentation and validation
Tutorials and guides
- Tutorials now use current replay-buffer storage, CUDA-graph, multi-agent reset, transform, and TRL interoperability patterns; their multiprocessing setup also works on Windows. The recurrent DQN material clarifies the required
TensorDictPrimerstate lifecycle, and the MicroDuck tutorial covers task libraries, reward registration, task sampling, per-task advantages, backends, gait control, and training. #4181 and #4219 by @vmoens, #4123 by @coder-jayp, and #4220 by @YeonwooSung.
API reference
- The value-objective reference documents PQN lambda returns through
DQNLosswithValueEstimators.TDLambda, including greedy next-action bootstrapping. #4233 by @YeonwooSung.
Examples, tests, benchmarks, and CI
- Examples have systematic CI execution, AsyncEnvPool has continuous dispatch and throughput benchmarks, and regression coverage was strengthened for noisy layers, SFT loss, reward-to-go values, tensor lambdas, chat-template round trips, optional RSSM scan dependencies, and repository-independent render test paths. #4152 and #4182 by @vmoens, #4163, #4168, #4170, and #4171 by @quinnarnold, #4235 by @YeonwooSung, and #4242 and #4243 by @coder-jayp.
-
Release maintenance refreshes the expert-iteration NLTK and Transformers dependencies, records the repository contribution contracts, and exercises MuJoCo-Torch from main alongside nightly PyTorch on macOS. #4199, #4201, #4207, and #4211.
-
Continuous benchmarks track async collection across merges, direct process-slot collection, benchmarked commit identities, and baseline collector and DreamerV3 measurements. Performance alerts no longer prevent result publication, and labelled merges can run the complete benchmark suite. #4287, #4289, #4292, #4293, #4298, and #4312 by @vmoens.
- The DreamerV3 process-inference benchmark selects the transport-owned environment exchange so both thread and process series remain runnable. #4320 by @vmoens.
- Regression coverage initializes MicroDuck fixtures deterministically, exercises prioritized-replay weights in DDPG, and preserves metadata in static inference batches. #4279 by @coder-jayp and #4288 and #4303 by @vmoens.
- Contributor guidance records target-device tensor and module construction as a repository contract. #4314 by @vmoens.
Upgrade notes
- Install matching release lines together:
pip install "torchrl==0.14.0" "tensordict>=0.14.2,<0.15.0". - Review code using
QValueModuleorQValueActorwith shaped categorical specs. Choose strict validation, automatic reshaping, or explicit opt-out rather than relying on the former warning-only default. - Replace
MultiCollector.postprocswithMultiCollector.postproc, and review the v0.15 checkpoint and collector-default warnings before the next release. - Pass an explicit
exchangeto multiprocessingAsyncEnvPoolinstances to keep queue transport or opt into automatic shared-memory selection before the default changes.
Full changelog
For the complete commit-by-commit history, see v0.13.3...v0.14.0.
Project stewardship
We are delighted to announce that @theap06 is joining @vmoens as a maintainer of TorchRL. His exceptional work across model-based reinforcement learning, replay buffers, checkpointing, and the library as a whole has made TorchRL substantially stronger. Thank you, @theap06, for the care, energy, and extraordinary amount of work you have brought to the project. We are looking forward to building and shipping what comes next together, alongside the entire TorchRL community.
Contributors
Thanks to @Agade09, @aswanth-07, @bsprenger, @coder-jayp, @discobot, @fallintoplace, @gtnv, @harryfrzz, @Iliamsou, @itwasabhi, @lin-erica, @mathieuorhan, @ParamThakkar123, @quinnarnold, @saputkin, @theap06, @unography, @vmoens, @xyf5432, @YeonwooSung, @younik, and @yurekami for the changes included in this release.
We especially welcome first-time contributors @Agade09, @aswanth-07, @discobot, @fallintoplace, @gtnv, @harryfrzz, @Iliamsou, @quinnarnold, @saputkin, @unography, @xyf5432, @YeonwooSung, and @yurekami.