| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| NVIDIA NeMo Speech 3.0 source code.tar.gz | 2026-08-06 | 57.2 MB | |
| NVIDIA NeMo Speech 3.0 source code.zip | 2026-08-06 | 59.2 MB | |
| README.md | 2026-08-06 | 20.3 kB | |
| Totals: 3 Items | 116.5 MB | 3 | |
NVIDIA NeMo Speech 3.0 Release Notes
NeMo Speech 3.0 is the first major release after the repo split and rename to NVIDIA-NeMo/Speech. The repo now focuses on ASR, TTS, audio processing, speaker tasks, and SpeechLM. Non-speech Framework, LLM, VLM, diffusion, export/deploy, evaluator components can now be found under separate repos under NVIDIA-NeMo organization.
While this release brings new major features, the central focus is addressing technical debt: removed 800k deprecated LOC, migration to uv package manager for cleaner installs, stronger model test coverage, reduced the number of dependencies, revamped documentation, lighter containers, and AGENTS.md + agentic skills.
Top-Level Changes and Highlights
Breaking Changes and Migration Notes
- LLM/VLM/diffusion/NLP/export/evaluator collections have been moved out to separate repos.
- Unsupported ASR/TTS models and old tutorials were removed. Use active examples under
examples/and docs underdocs/source/. uv syncreproduces the supported NeMo Speech stack and may replace Python/PyTorch/CUDA inside.venv. To keep your own stack, install PyTorch first, then useuv piporpip.- SpeechLM2 compiled acceleration is optional. Use the supported Docker build for TE/FlashAttention/Mamba/grouped-GEMM/DeepEP performance paths.
SpeechLM2 and NeMo Automodel Backend
SpeechLM2 now supports training SpeechLM models with NeMo Automodel through SALMAutomodel, targeting both dense and MoE LLM backbones such as Nemotron 3.
SALMAutomodeladds an Automodel-backed SALM path with native LoRA, deferredconfigure_model()initialization, shard-aware distributed loading, and Automodel-owned mesh creation (#15447).AutomodelParallelStrategysupports FSDP2, HSDP, TP, CP, and EP, letting SpeechLM training combine sharded data parallelism, tensor/model partitioning, long-context sequence sharding, and MoE expert routing (#15447, [#15648], [#15679], [#15773]).- Optimized backend support now covers Transformer Engine, FlashAttention, Mamba/state-space kernels, grouped GEMM for MoE experts, and DeepEP for expert-parallel all-to-all communication (#15737, [#15758], [#15679]).
- THD packed-sequence training reduces padding waste in variable-length speech batches and is the supported path for CP long-audio training (#15679).
- Long-audio support was expanded with encoder chunking, activation-checkpointing controls, CP-safe data handling, and fixes around dtype/checkpoint restoration (#15648, [#15679], [#15686], [#15716]).
- Export and serving improved with vLLM integration, HF/vLLM-ready checkpoint export, backbone-native chat templates, offline HF export, and buffered SALM inference (#15364, [#15520], [#15623], [#15736]).
- Data format support now includes ShareGPT conversations, indexed ShareGPT JSONL, and WebDataset-backed ShareGPT conversations (#15316, [#15410]).
Nemotron VoiceChat Training Modules
NeMo Speech 3.0 adds the train/eval modules behind Nemotron VoiceChat.
- Added Nemotron VoiceChat speech-decoder modules for the response-generation side of the duplex stack (#15066).
- Added DuplexSTT training and inference for generating agent text from user speech plus text context (#15092).
- Added Duplex/EAR-TTS speech-generation pieces used by VoiceChat, including MagpieTTS decoder integration and semantic-codec training support (#15277, [#15524]).
- Added the joint
NemotronVoiceChatSTT + TTS class for validation, offline speech-to-speech inference, and export workflows (#15456). - Added formatter/reproducibility work for the VoiceChat speech decoder, with faster training and half-precision inference support (#15583).
- Expanded Voice Agent examples for Nemotron Nano v2/v3, MagpieTTS, tool calling, audio logging, and README guidance (#14325, [#14704], [#15243], [#15269], [#15318], [#15547], [#15642]).
- Hardened Voice Agent behavior around default Parakeet EOU STT, text aggregation, end-of-bot handling, RTVI messages, empty tools, and logging (#14951, [#15068], [#15069], [#15634]).
Note: NemotronVoiceChat class is inference/eval only. Train DuplexSTTModel and DuplexEARTTS/speech-decoder modules separately.
ASR
- Unified prompt-model support now covers multilingual ASR and streaming inference (#15666).
- Streaming ASR gained Unified RNNT inference, batched streaming beam search, simulated chunked transducer decoding, streaming speech translation, and Canary streaming policies (#15522, [#15753], [#15517], [#15132], [#14765]).
- Transducer decoding is faster and more complete: 2.4x speedup with confidence, buffered/cache-aware confidence, fewer redundant decoding passes, lower streaming memory, and WER fixes (#15315, [#15765], [#15301], [#15148]).
- Phrase boosting expanded across transducers, cache-aware RNN-T, CTC/RNN-T/TDT GPU phrase boosting, and cache-aware customization (#15125, [#15344], [#14277], [#14757], [#14800]).
- New and updated model families include prompt Parakeet Hybrid RNNT/CTC, ASR EOU models, Streaming Sortformer, Canary2 with NFA, and ASR Transformer/Conformer-Transformer encoders (#14561, [#14740], [#14627], [#14121], [#15661], [#15703]).
- Diarization/VAD removed the Pyannote dependency and gained more flexible input handling (#15632, [#15184]).
TTS
- MagpieTTS received a major refresh: decoder model, refactor, longform inference, unified longform/standard paths, and updated model internals (#15277, [#15504], [#15210], [#15241], [#15477], [#15375], [#15031]).
- Added CFG distillation, local transformer CFG distillation, and MagpieTTS MoE support (#15568, [#15748], [#15370]).
- Expanded language/text support: Hindi, Japanese, Arabic char tokenizer, Japanese-English katakana, hi-IN/ko-KR/pt-BR IPA, and Japanese G2P accent support (#15320, [#15248], [#15614], [#15567], [#15170]).
- Improved EasyMagpie zero-shot disabling and speaker-encoder behavior (#15639, [#15503], [#15564]).
- Codec and evaluation work added semantic-codec training, codec conversion/bandwidth extension, HF audio-codec loading, Frechet Codec Distance, parallel scoring, comparison reports, and deterministic multi-GPU eval (#15524, [#15191], [#15172], [#15223], [#15417], [#15621], [#15427]).
- Inference paths now support reference-free inference, dataset selection, clearer inference config, and fp16 Whisper loading (#15213, [#15212], [#15254], [#15680]).
Audio
- Added and fixed Conformer U-Net speech-enhancement models (#14442, [#14626]).
- Added a data-prediction objective for flow-matching speech enhancement (#14749).
- Added streaming mode to
SpectrogramToAudio(#14524). - Fixed BNR 2.0 inference alignment with padded input signals (#15388).
- Reduced duplication and applied small fixes in the audio collection (#15587).
DataLoader and Speech Data Tools
- Lhotse support expanded to Parquet/Arrow embedded audio, AIS batch loading, AIS-hosted input configs, non-tarred S3 audio, temperature re-weighting, randomized tarred shard slicing, and in-manifest channel selection (#15303, [#15102], [#15538], [#14891], [#15732], [#15200], [#14558], [#14586]).
- Audio-codec data loading gained Lhotse training format support plus follow-up fixes (#15622, [#15742]).
- Online data augmentation now covers clipping, lowpass, lossy codec augmentation, and saving augmented audio from Lhotse samplers/dataloaders (#14809, [#14808]).
- Speech Data Explorer gained S3 reading, comparison mode, security fixes, and tutorial updates (#15500, [#15137]).
Docs
- Documentation was rewritten around speech scope, uv installs, ASR/TTS/SpeechLM2 guides, and cleaner navigation.
Repo Rename and 3.0 Release
- Project identity is now NVIDIA NeMo Speech in
NVIDIA-NeMo/Speech(#15783, [#15788]). - README/docs now describe the speech-focused repo, current status, and nightly docs (#15127, [#15217], [#15602], [#15766]).
Technical Debt Reduction
- Removed LLM, VLM, diffusion, NLP, multimodal, ExportDeploy, evaluator, Automodel, and language-modeling paths (#14095, [#14192], [#14617], [#14677], [#14790], [#14934], [#14964], [#15033], [#15044], [#15357], [#15378]).
- Cleaned unsupported ASR/TTS modules, old tutorials, old buffered CTC scripts, and deprecated SpeechLM1 scripts (#14660, [#15061], [#15368], [#15507], [#15655]).
- Removed unsafe pickle paths and hardened target/export handling (#14540, [#15065], [#15232], [#15266], [#15288], [#15629]).
Installation, Packaging, and Containers
- Switched to a uv-first source install:
uv sync --extra all --extra cu13(#15769). - Added clear bring-your-own Python/PyTorch/CUDA guidance via
uv piporpip, addressing common install feedback (#15769). - Routed optional SpeechLM2 compiled deps through uv-compatible sources and container builds: Transformer Engine, FlashAttention, Mamba, grouped GEMM, DeepEP (#15737, [#15758]).
- Refreshed Docker around official PyTorch bases, lighter staged builds, CUDA 12/13 support, and CVE dependency updates (#15638, [#15747], [#15756], [#15761]).
- Updated CI/publish flows to install from
uv.lockand simplify release builds (#15659, [#15668], [#15685], [#15697]).
Tests and Release Hardening
- Added functional init/train-step/inference tests for every supported released model (#15433).
- Added golden-value MoE dispatch tests for grouped GEMM/expert routing (#15698).
- Split L0 tests into ASR GPU/CPU and SpeechLM2 buckets (#15654).
- Hardened streaming ASR, Sortformer, MagpieTTS, codec, and TTS L2/e2e tests (#14416, [#14417], [#14435], [#14823], [#15272], [#15508], [#15584], [#15607]).
- Refactored release CI and moved to AWS ephemeral runners (#15620, [#15668], [#15718]).
Detailed PR Breakdown
Release Identity and Docs
- [#15788] releases
r3.0.0; [#15783] bumps the next release to 3.0. - [#15127] renames docs/title to NVIDIA NeMo Speech; [#15217] and [#15602] update repo status/release messaging.
- [#15363] and [#15460] restructure docs; [#15542] refactors ASR docs; [#15647] refactors diarization docs; [#15745] improves docs visuals.
- [#15769] rewrites install docs for uv, bring-your-own stacks, CUDA extras, and pip fallback.
- [#15773] documents SpeechLM parallelism; [#15738] adds prompt RNNT docs; [#15546]/#15302 add MagpieTTS docs; [#15526] adds streaming ASR inference docs.
- [#15744] adds community featured models; [#15766] updates the release README.
Technical-Debt Cleanup
- [#15378] removes deprecated collections; [#15363] starts docs-side transformation.
- [#14095], [#14677], [#14934], [#14617], and [#14192] remove NeMo1, multimodal/vision, NLP, and language-modeling code.
- [#15357], [#14790], [#15033], [#15044], and [#14964] remove deprecated LLM/VLM/diffusion tutorials, export-deploy, old Automodel, and evaluator code.
- [#15507] cleans old ASR modules; [#15368] removes unsupported TTS models/docs; [#14660] removes old TTS tutorials; [#15061] removes old buffered CTC; [#15655] removes SpeechLM1 scripts.
- [#15147] updates TRT references; [#15146] removes old checkpoint save support; [#15211] removes torchaudio usage.
- [#15687] applies black formatting; [#15656] makes formatting CI check-only with pre-commit hooks.
Install, Packaging, Docker, Dependencies
- [#15769] adds uv-first install plus pip/uv-pip fallback for existing PyTorch/CUDA stacks.
- [#15734] removes stale uv-base deps; [#15758] moves DeepEP/TE git refs to uv sources; [#15697] installs CI deps from
uv.lock. - [#15638] switches Docker to official PyTorch base; [#15747] adds CUDA 12 builds; [#15761] moves code copy to the final Docker stage.
- [#15737] builds Automodel compiled deps in CI image; [#15756] updates CVE-sensitive deps and Dockerfile.
- [#15659] unblocks TestPyPI publish; [#15668] refactors release workflows; [#15685] removes obsolete build workflows.
- [#15314] sets
torch >= 2.6.0; [#15353] documentsTORCH_FORCE_NO_WEIGHTS_ONLY_LOAD; [#15506]/#15018 fix CUDA Python/numba-cuda handling. - [#15630] removes
diskcache; [#15502] removes protobuf; [#15183] removes CUDA bindings pins; [#15438] relaxeskaldialign.
Security and Safety
- [#15288] fixes command injection and insecure permissions.
- [#14540] hardens target resolution.
- [#15065] guards
trust_remote_code. - [#15266], [#15232], and [#15629] remove or block unsafe pickle paths.
- [#15636] adds
SECURITY.md. - [#15500] adds SDE S3/comparison updates plus a security fix.
Tests, CI, Release Reliability
- [#15433] adds init/train-step/inference functional tests for released models.
- [#15698] adds golden MoE dispatch tests.
- [#15654] splits L0 ASR/SpeechLM2 test buckets; [#15651] raises GPU speech test timeouts.
- [#14823], [#14953], and [#14184] harden streaming/cache-aware transducer and CUDA-graph tests.
- [#14416], [#14417], and [#14435] complete Streaming Sortformer fixes/tests.
- [#15508], [#15584], [#15272], and [#15607] harden TTS/MagpieTTS/codec tests.
- [#15620], [#15718], [#15484], [#15519], and [#15501] improve CI runners, branch rules, triggering, and e2e layout.
- [#15613] adds local dev/PR babysitting support; [#15612] adds distributed-log debug support.
SpeechLM2, SALMAutomodel, Scaling, vLLM
- [#15447] adds SALM + NeMo Automodel for Nemotron Nano V3, including native LoRA and Automodel-managed distributed setup.
- [#15648] expands
SALMAutomodelwith long-context support, encoder chunking, activation-checkpointing controls, and stability fixes. - [#15679] adds THD packed sequences, context parallelism, CP-safe long-audio training, and related packed-input handling.
- [#15773] documents FSDP2, HSDP, TP, PP, CP, and EP strategy usage for SpeechLM training.
- [#15737] and [#15758] wire compiled Automodel deps, DeepEP, and Transformer Engine into reproducible uv/container builds.
- [#15520] adds vLLM support for NeMo SpeechLM, including a registered SpeechLM model/config path.
- [#15623] exports vLLM-ready SpeechLM checkpoints with backbone-native
chat_templatesupport. - [#15716] adds encoder input chunking for SALM vLLM inference.
- [#15736] adds offline HF export for SpeechLM2 checkpoints.
- [#15686] preserves perception checkpoint dtype; [#15570] fixes Qwen3 SALM LoRA initialization.
- [#15316] adds TDT decoder input, ShareGPT format support, and related SALM improvements.
- [#15410] adds indexed ShareGPT JSONL and WebDataset formats.
- [#15364] adds buffered SALM inference; [#15478] updates SpeechLM2 docs and dataloader resampling; [#15281] adds multimodal offset-key support.
Nemotron VoiceChat, Duplex Models, Voice Agent
- [#15066] implements the Nemotron VoiceChat speech decoder for the response-generation side of duplex speech-to-speech.
- [#15092] implements DuplexSTT for speech/text-conditioned agent text generation.
- [#15456] adds the joint Nemotron VoiceChat STT + TTS class for validation, offline inference, and export.
- [#15583] improves VoiceChat decoder reproducibility, training speed, formatting, and fp16 inference.
- [#15277] adds the MagpieTTS decoder model used by the speech-generation stack.
- [#15524] adds semantic-codec training support.
- [#14325] adds NeMo Voice Agent as the app/integration layer.
- [#14704] adds Nemotron Nano v2 support; [#15318] adds Nemotron Nano v3 and MagpieTTS improvements.
- [#15069] sets Parakeet EOU as default STT for the voice agent.
- [#14951] fixes text aggregation, end-of-bot handling, and logging; [#15068] fixes missing RTVI bot messages.
- [#15243] and [#15269] improve tool-calling examples, tool-calling behavior, and logging UX.
- [#15634] fixes empty tools; [#15642] updates Voice Agent docs.
ASR, Streaming, Prompt Models, Transducers
- [#15666] adds unified prompt-model support for multilingual ASR and streaming; [#14561] adds Parakeet hybrid prompt support; [#15738] adds docs.
- [#15522] adds Unified RNNT streaming inference.
- [#15753] adds batched streaming beam search; [#15315] speeds up transducer decoding 2.4x with confidence; [#15765] adds buffered/cache-aware confidence.
- [#15517] adds simulated chunked transducer decoding; [#15301] removes double decoding; [#15148] fixes streaming transducer memory/WER.
- [#15125], [#15344], [#14277], [#14800], and [#14757] expand phrase boosting and cache-aware customization.
- [#15132] adds streaming speech translation; [#14765] adds Canary Wait-K/AlignAtt; [#14766] adds streaming timestamps.
- [#15268] adds Canary timestamp-model restore; [#15291] fixes tensor-input timestamps.
- [#14740], [#15462], [#15497], and [#15493] add EOU models/metrics and fixes.
- [#14627], [#14388], [#14416], [#14417], and [#14435] cover Streaming Sortformer train/inference/docs/tests.
- [#14905], [#14932], and [#15025] update Multi-Talker Parakeet streaming docs/notebooks; [#14121] adds Canary2 with NFA; [#15317] improves short-audio Canary.
- [#15632] removes Pyannote from diarization/VAD; [#15184] improves diarization inputs; [#15573] aligns cpWER with meeteval.
- [#15605] removes tail-margin algorithms; [#15591] removes PnC flag; [#15576] fixes ASR inference pipeline bugs.
- [#15661] adds ASR Transformer encoder; [#15703] adds Conformer I/O-styled Transformer encoder.
TTS, MagpieTTS, Codecs, Text
- [#15277] adds MagpieTTS decoder; [#15504] refactors MagpieTTS; [#15031] updates the model.
- [#15210], [#15241], [#15477], and [#15375] add/unify longform MagpieTTS inference.
- [#15568] and [#15748] add CFG distillation; [#15370] adds MagpieTTS MoE.
- [#15320], [#15248], [#15567], [#15614], and [#15170] add Hindi/Japanese/IPA/Arabic/Katakana/G2P support.
- [#15639], [#15503], and [#15564] adjust EasyMagpie zero-shot disabling.
- [#15191], [#15172], [#15622], and [#15742] improve codec conversion, bandwidth extension, HF codec loading, and Lhotse codec data.
- [#15524] adds semantic-codec training; [#15223] adds Frechet Codec Distance; [#15417] adds parallel scoring; [#15621] adds TTS comparison reports.
- [#15348] and [#15427] improve multi-validation and deterministic multi-GPU eval.
- [#15213], [#15212], [#15254], and [#15680] improve reference-free inference, dataset selection, inference config, and fp16 Whisper loading.
Audio
- [#14442] adds a Conformer U-Net model for speech enhancement.
- [#14626] fixes the Conformer U-Net implementation.
- [#14749] adds a data-prediction objective for flow-matching speech-enhancement models.
- [#14524] adds streaming mode to
SpectrogramToAudio. - [#15388] fixes BNR 2.0 inference alignment with padded input signals.
- [#15587] reduces duplication in the audio collection and applies small fixes.
DataLoader and Speech Data Tools
- [#15303] adds Parquet/Arrow embedded-audio support via Lhotse.
- [#15102] and [#15538] add AIS batch loading for ASR audio processing and
LhotseSpeechToTextBpeDataset. - [#14891] supports
input_cfg.yamldirectly from AIS buckets. - [#15732] fixes non-tarred S3 audio in
LazyNeMoIterator. - [#15200] adds on-the-fly temperature re-weighting for Lhotse datasets.
- [#14558] adds randomized shard slicing for tarred data.
- [#14586] adds in-manifest channel selection for multichannel recordings through Lhotse.
- [#15316] and [#15410] add ShareGPT, indexed ShareGPT JSONL, and WebDataset data formats for SpeechLM2.
- [#14294] adds ShareGPT data and a test loader.
- [#15622] adds audio-codec Lhotse training format support; [#15742] updates and fixes audio-codec Lhotse loading.
- [#14809] adds clipping, lowpass, and lossy-codec online augmentations; [#14808] adds augmented-audio saving.
- [#15500] updates Speech Data Explorer with S3 read support, comparison mode, and security fixes; [#15137] updates the SDE tutorial.
Smaller Fixes
- [#15437] adds
.nemotorchaudio-preprocessor checkpoint migration. - [#15461] auto-detects
use_bucketingand validates batch size. - [#15325] fixes TDT fusion scores; [#15322] moves transducer fusion models to the base class.
- [#15245] speeds
.nemotar extraction/config processing. - [#15409] handles an exp-manager timer race.
- [#15238] uses
PurePosixPath. - [#15444] replaces
assertwithValueErrorinEncDecMultiTaskModel. - [#15688] fixes validation audio logging; [#15699] adds
UTMOSv2Calculator.verbose.