Download Latest Version NVIDIA NeMo Speech 3.0 source code.zip (59.2 MB) Google Add to Preferred Sources
Home / v3.0.0
Name Modified Size InfoDownloads / Week
Parent folder
NVIDIA NeMo Speech 3.0 source code.tar.gz 2026-08-06 57.2 MB
NVIDIA NeMo Speech 3.0 source code.zip 2026-08-06 59.2 MB
README.md 2026-08-06 20.3 kB
Totals: 3 Items   116.5 MB 3

NVIDIA NeMo Speech 3.0 Release Notes

NeMo Speech 3.0 is the first major release after the repo split and rename to NVIDIA-NeMo/Speech. The repo now focuses on ASR, TTS, audio processing, speaker tasks, and SpeechLM. Non-speech Framework, LLM, VLM, diffusion, export/deploy, evaluator components can now be found under separate repos under NVIDIA-NeMo organization.

While this release brings new major features, the central focus is addressing technical debt: removed 800k deprecated LOC, migration to uv package manager for cleaner installs, stronger model test coverage, reduced the number of dependencies, revamped documentation, lighter containers, and AGENTS.md + agentic skills.

Top-Level Changes and Highlights

Breaking Changes and Migration Notes

  • LLM/VLM/diffusion/NLP/export/evaluator collections have been moved out to separate repos.
  • Unsupported ASR/TTS models and old tutorials were removed. Use active examples under examples/ and docs under docs/source/.
  • uv sync reproduces the supported NeMo Speech stack and may replace Python/PyTorch/CUDA inside .venv. To keep your own stack, install PyTorch first, then use uv pip or pip.
  • SpeechLM2 compiled acceleration is optional. Use the supported Docker build for TE/FlashAttention/Mamba/grouped-GEMM/DeepEP performance paths.

SpeechLM2 and NeMo Automodel Backend

SpeechLM2 now supports training SpeechLM models with NeMo Automodel through SALMAutomodel, targeting both dense and MoE LLM backbones such as Nemotron 3.

  • SALMAutomodel adds an Automodel-backed SALM path with native LoRA, deferred configure_model() initialization, shard-aware distributed loading, and Automodel-owned mesh creation (#15447).
  • AutomodelParallelStrategy supports FSDP2, HSDP, TP, CP, and EP, letting SpeechLM training combine sharded data parallelism, tensor/model partitioning, long-context sequence sharding, and MoE expert routing (#15447, [#15648], [#15679], [#15773]).
  • Optimized backend support now covers Transformer Engine, FlashAttention, Mamba/state-space kernels, grouped GEMM for MoE experts, and DeepEP for expert-parallel all-to-all communication (#15737, [#15758], [#15679]).
  • THD packed-sequence training reduces padding waste in variable-length speech batches and is the supported path for CP long-audio training (#15679).
  • Long-audio support was expanded with encoder chunking, activation-checkpointing controls, CP-safe data handling, and fixes around dtype/checkpoint restoration (#15648, [#15679], [#15686], [#15716]).
  • Export and serving improved with vLLM integration, HF/vLLM-ready checkpoint export, backbone-native chat templates, offline HF export, and buffered SALM inference (#15364, [#15520], [#15623], [#15736]).
  • Data format support now includes ShareGPT conversations, indexed ShareGPT JSONL, and WebDataset-backed ShareGPT conversations (#15316, [#15410]).

Nemotron VoiceChat Training Modules

NeMo Speech 3.0 adds the train/eval modules behind Nemotron VoiceChat.

  • Added Nemotron VoiceChat speech-decoder modules for the response-generation side of the duplex stack (#15066).
  • Added DuplexSTT training and inference for generating agent text from user speech plus text context (#15092).
  • Added Duplex/EAR-TTS speech-generation pieces used by VoiceChat, including MagpieTTS decoder integration and semantic-codec training support (#15277, [#15524]).
  • Added the joint NemotronVoiceChat STT + TTS class for validation, offline speech-to-speech inference, and export workflows (#15456).
  • Added formatter/reproducibility work for the VoiceChat speech decoder, with faster training and half-precision inference support (#15583).
  • Expanded Voice Agent examples for Nemotron Nano v2/v3, MagpieTTS, tool calling, audio logging, and README guidance (#14325, [#14704], [#15243], [#15269], [#15318], [#15547], [#15642]).
  • Hardened Voice Agent behavior around default Parakeet EOU STT, text aggregation, end-of-bot handling, RTVI messages, empty tools, and logging (#14951, [#15068], [#15069], [#15634]).

Note: NemotronVoiceChat class is inference/eval only. Train DuplexSTTModel and DuplexEARTTS/speech-decoder modules separately.

ASR

  • Unified prompt-model support now covers multilingual ASR and streaming inference (#15666).
  • Streaming ASR gained Unified RNNT inference, batched streaming beam search, simulated chunked transducer decoding, streaming speech translation, and Canary streaming policies (#15522, [#15753], [#15517], [#15132], [#14765]).
  • Transducer decoding is faster and more complete: 2.4x speedup with confidence, buffered/cache-aware confidence, fewer redundant decoding passes, lower streaming memory, and WER fixes (#15315, [#15765], [#15301], [#15148]).
  • Phrase boosting expanded across transducers, cache-aware RNN-T, CTC/RNN-T/TDT GPU phrase boosting, and cache-aware customization (#15125, [#15344], [#14277], [#14757], [#14800]).
  • New and updated model families include prompt Parakeet Hybrid RNNT/CTC, ASR EOU models, Streaming Sortformer, Canary2 with NFA, and ASR Transformer/Conformer-Transformer encoders (#14561, [#14740], [#14627], [#14121], [#15661], [#15703]).
  • Diarization/VAD removed the Pyannote dependency and gained more flexible input handling (#15632, [#15184]).

TTS

  • MagpieTTS received a major refresh: decoder model, refactor, longform inference, unified longform/standard paths, and updated model internals (#15277, [#15504], [#15210], [#15241], [#15477], [#15375], [#15031]).
  • Added CFG distillation, local transformer CFG distillation, and MagpieTTS MoE support (#15568, [#15748], [#15370]).
  • Expanded language/text support: Hindi, Japanese, Arabic char tokenizer, Japanese-English katakana, hi-IN/ko-KR/pt-BR IPA, and Japanese G2P accent support (#15320, [#15248], [#15614], [#15567], [#15170]).
  • Improved EasyMagpie zero-shot disabling and speaker-encoder behavior (#15639, [#15503], [#15564]).
  • Codec and evaluation work added semantic-codec training, codec conversion/bandwidth extension, HF audio-codec loading, Frechet Codec Distance, parallel scoring, comparison reports, and deterministic multi-GPU eval (#15524, [#15191], [#15172], [#15223], [#15417], [#15621], [#15427]).
  • Inference paths now support reference-free inference, dataset selection, clearer inference config, and fp16 Whisper loading (#15213, [#15212], [#15254], [#15680]).

Audio

  • Added and fixed Conformer U-Net speech-enhancement models (#14442, [#14626]).
  • Added a data-prediction objective for flow-matching speech enhancement (#14749).
  • Added streaming mode to SpectrogramToAudio (#14524).
  • Fixed BNR 2.0 inference alignment with padded input signals (#15388).
  • Reduced duplication and applied small fixes in the audio collection (#15587).

DataLoader and Speech Data Tools

  • Lhotse support expanded to Parquet/Arrow embedded audio, AIS batch loading, AIS-hosted input configs, non-tarred S3 audio, temperature re-weighting, randomized tarred shard slicing, and in-manifest channel selection (#15303, [#15102], [#15538], [#14891], [#15732], [#15200], [#14558], [#14586]).
  • Audio-codec data loading gained Lhotse training format support plus follow-up fixes (#15622, [#15742]).
  • Online data augmentation now covers clipping, lowpass, lossy codec augmentation, and saving augmented audio from Lhotse samplers/dataloaders (#14809, [#14808]).
  • Speech Data Explorer gained S3 reading, comparison mode, security fixes, and tutorial updates (#15500, [#15137]).

Docs

  • Documentation was rewritten around speech scope, uv installs, ASR/TTS/SpeechLM2 guides, and cleaner navigation.

Repo Rename and 3.0 Release

  • Project identity is now NVIDIA NeMo Speech in NVIDIA-NeMo/Speech (#15783, [#15788]).
  • README/docs now describe the speech-focused repo, current status, and nightly docs (#15127, [#15217], [#15602], [#15766]).

Technical Debt Reduction

  • Removed LLM, VLM, diffusion, NLP, multimodal, ExportDeploy, evaluator, Automodel, and language-modeling paths (#14095, [#14192], [#14617], [#14677], [#14790], [#14934], [#14964], [#15033], [#15044], [#15357], [#15378]).
  • Cleaned unsupported ASR/TTS modules, old tutorials, old buffered CTC scripts, and deprecated SpeechLM1 scripts (#14660, [#15061], [#15368], [#15507], [#15655]).
  • Removed unsafe pickle paths and hardened target/export handling (#14540, [#15065], [#15232], [#15266], [#15288], [#15629]).

Installation, Packaging, and Containers

  • Switched to a uv-first source install: uv sync --extra all --extra cu13 (#15769).
  • Added clear bring-your-own Python/PyTorch/CUDA guidance via uv pip or pip, addressing common install feedback (#15769).
  • Routed optional SpeechLM2 compiled deps through uv-compatible sources and container builds: Transformer Engine, FlashAttention, Mamba, grouped GEMM, DeepEP (#15737, [#15758]).
  • Refreshed Docker around official PyTorch bases, lighter staged builds, CUDA 12/13 support, and CVE dependency updates (#15638, [#15747], [#15756], [#15761]).
  • Updated CI/publish flows to install from uv.lock and simplify release builds (#15659, [#15668], [#15685], [#15697]).

Tests and Release Hardening

  • Added functional init/train-step/inference tests for every supported released model (#15433).
  • Added golden-value MoE dispatch tests for grouped GEMM/expert routing (#15698).
  • Split L0 tests into ASR GPU/CPU and SpeechLM2 buckets (#15654).
  • Hardened streaming ASR, Sortformer, MagpieTTS, codec, and TTS L2/e2e tests (#14416, [#14417], [#14435], [#14823], [#15272], [#15508], [#15584], [#15607]).
  • Refactored release CI and moved to AWS ephemeral runners (#15620, [#15668], [#15718]).

Detailed PR Breakdown

Release Identity and Docs

  • [#15788] releases r3.0.0; [#15783] bumps the next release to 3.0.
  • [#15127] renames docs/title to NVIDIA NeMo Speech; [#15217] and [#15602] update repo status/release messaging.
  • [#15363] and [#15460] restructure docs; [#15542] refactors ASR docs; [#15647] refactors diarization docs; [#15745] improves docs visuals.
  • [#15769] rewrites install docs for uv, bring-your-own stacks, CUDA extras, and pip fallback.
  • [#15773] documents SpeechLM parallelism; [#15738] adds prompt RNNT docs; [#15546]/#15302 add MagpieTTS docs; [#15526] adds streaming ASR inference docs.
  • [#15744] adds community featured models; [#15766] updates the release README.

Technical-Debt Cleanup

  • [#15378] removes deprecated collections; [#15363] starts docs-side transformation.
  • [#14095], [#14677], [#14934], [#14617], and [#14192] remove NeMo1, multimodal/vision, NLP, and language-modeling code.
  • [#15357], [#14790], [#15033], [#15044], and [#14964] remove deprecated LLM/VLM/diffusion tutorials, export-deploy, old Automodel, and evaluator code.
  • [#15507] cleans old ASR modules; [#15368] removes unsupported TTS models/docs; [#14660] removes old TTS tutorials; [#15061] removes old buffered CTC; [#15655] removes SpeechLM1 scripts.
  • [#15147] updates TRT references; [#15146] removes old checkpoint save support; [#15211] removes torchaudio usage.
  • [#15687] applies black formatting; [#15656] makes formatting CI check-only with pre-commit hooks.

Install, Packaging, Docker, Dependencies

  • [#15769] adds uv-first install plus pip/uv-pip fallback for existing PyTorch/CUDA stacks.
  • [#15734] removes stale uv-base deps; [#15758] moves DeepEP/TE git refs to uv sources; [#15697] installs CI deps from uv.lock.
  • [#15638] switches Docker to official PyTorch base; [#15747] adds CUDA 12 builds; [#15761] moves code copy to the final Docker stage.
  • [#15737] builds Automodel compiled deps in CI image; [#15756] updates CVE-sensitive deps and Dockerfile.
  • [#15659] unblocks TestPyPI publish; [#15668] refactors release workflows; [#15685] removes obsolete build workflows.
  • [#15314] sets torch >= 2.6.0; [#15353] documents TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD; [#15506]/#15018 fix CUDA Python/numba-cuda handling.
  • [#15630] removes diskcache; [#15502] removes protobuf; [#15183] removes CUDA bindings pins; [#15438] relaxes kaldialign.

Security and Safety

  • [#15288] fixes command injection and insecure permissions.
  • [#14540] hardens target resolution.
  • [#15065] guards trust_remote_code.
  • [#15266], [#15232], and [#15629] remove or block unsafe pickle paths.
  • [#15636] adds SECURITY.md.
  • [#15500] adds SDE S3/comparison updates plus a security fix.

Tests, CI, Release Reliability

  • [#15433] adds init/train-step/inference functional tests for released models.
  • [#15698] adds golden MoE dispatch tests.
  • [#15654] splits L0 ASR/SpeechLM2 test buckets; [#15651] raises GPU speech test timeouts.
  • [#14823], [#14953], and [#14184] harden streaming/cache-aware transducer and CUDA-graph tests.
  • [#14416], [#14417], and [#14435] complete Streaming Sortformer fixes/tests.
  • [#15508], [#15584], [#15272], and [#15607] harden TTS/MagpieTTS/codec tests.
  • [#15620], [#15718], [#15484], [#15519], and [#15501] improve CI runners, branch rules, triggering, and e2e layout.
  • [#15613] adds local dev/PR babysitting support; [#15612] adds distributed-log debug support.

SpeechLM2, SALMAutomodel, Scaling, vLLM

  • [#15447] adds SALM + NeMo Automodel for Nemotron Nano V3, including native LoRA and Automodel-managed distributed setup.
  • [#15648] expands SALMAutomodel with long-context support, encoder chunking, activation-checkpointing controls, and stability fixes.
  • [#15679] adds THD packed sequences, context parallelism, CP-safe long-audio training, and related packed-input handling.
  • [#15773] documents FSDP2, HSDP, TP, PP, CP, and EP strategy usage for SpeechLM training.
  • [#15737] and [#15758] wire compiled Automodel deps, DeepEP, and Transformer Engine into reproducible uv/container builds.
  • [#15520] adds vLLM support for NeMo SpeechLM, including a registered SpeechLM model/config path.
  • [#15623] exports vLLM-ready SpeechLM checkpoints with backbone-native chat_template support.
  • [#15716] adds encoder input chunking for SALM vLLM inference.
  • [#15736] adds offline HF export for SpeechLM2 checkpoints.
  • [#15686] preserves perception checkpoint dtype; [#15570] fixes Qwen3 SALM LoRA initialization.
  • [#15316] adds TDT decoder input, ShareGPT format support, and related SALM improvements.
  • [#15410] adds indexed ShareGPT JSONL and WebDataset formats.
  • [#15364] adds buffered SALM inference; [#15478] updates SpeechLM2 docs and dataloader resampling; [#15281] adds multimodal offset-key support.

Nemotron VoiceChat, Duplex Models, Voice Agent

  • [#15066] implements the Nemotron VoiceChat speech decoder for the response-generation side of duplex speech-to-speech.
  • [#15092] implements DuplexSTT for speech/text-conditioned agent text generation.
  • [#15456] adds the joint Nemotron VoiceChat STT + TTS class for validation, offline inference, and export.
  • [#15583] improves VoiceChat decoder reproducibility, training speed, formatting, and fp16 inference.
  • [#15277] adds the MagpieTTS decoder model used by the speech-generation stack.
  • [#15524] adds semantic-codec training support.
  • [#14325] adds NeMo Voice Agent as the app/integration layer.
  • [#14704] adds Nemotron Nano v2 support; [#15318] adds Nemotron Nano v3 and MagpieTTS improvements.
  • [#15069] sets Parakeet EOU as default STT for the voice agent.
  • [#14951] fixes text aggregation, end-of-bot handling, and logging; [#15068] fixes missing RTVI bot messages.
  • [#15243] and [#15269] improve tool-calling examples, tool-calling behavior, and logging UX.
  • [#15634] fixes empty tools; [#15642] updates Voice Agent docs.

ASR, Streaming, Prompt Models, Transducers

  • [#15666] adds unified prompt-model support for multilingual ASR and streaming; [#14561] adds Parakeet hybrid prompt support; [#15738] adds docs.
  • [#15522] adds Unified RNNT streaming inference.
  • [#15753] adds batched streaming beam search; [#15315] speeds up transducer decoding 2.4x with confidence; [#15765] adds buffered/cache-aware confidence.
  • [#15517] adds simulated chunked transducer decoding; [#15301] removes double decoding; [#15148] fixes streaming transducer memory/WER.
  • [#15125], [#15344], [#14277], [#14800], and [#14757] expand phrase boosting and cache-aware customization.
  • [#15132] adds streaming speech translation; [#14765] adds Canary Wait-K/AlignAtt; [#14766] adds streaming timestamps.
  • [#15268] adds Canary timestamp-model restore; [#15291] fixes tensor-input timestamps.
  • [#14740], [#15462], [#15497], and [#15493] add EOU models/metrics and fixes.
  • [#14627], [#14388], [#14416], [#14417], and [#14435] cover Streaming Sortformer train/inference/docs/tests.
  • [#14905], [#14932], and [#15025] update Multi-Talker Parakeet streaming docs/notebooks; [#14121] adds Canary2 with NFA; [#15317] improves short-audio Canary.
  • [#15632] removes Pyannote from diarization/VAD; [#15184] improves diarization inputs; [#15573] aligns cpWER with meeteval.
  • [#15605] removes tail-margin algorithms; [#15591] removes PnC flag; [#15576] fixes ASR inference pipeline bugs.
  • [#15661] adds ASR Transformer encoder; [#15703] adds Conformer I/O-styled Transformer encoder.

TTS, MagpieTTS, Codecs, Text

  • [#15277] adds MagpieTTS decoder; [#15504] refactors MagpieTTS; [#15031] updates the model.
  • [#15210], [#15241], [#15477], and [#15375] add/unify longform MagpieTTS inference.
  • [#15568] and [#15748] add CFG distillation; [#15370] adds MagpieTTS MoE.
  • [#15320], [#15248], [#15567], [#15614], and [#15170] add Hindi/Japanese/IPA/Arabic/Katakana/G2P support.
  • [#15639], [#15503], and [#15564] adjust EasyMagpie zero-shot disabling.
  • [#15191], [#15172], [#15622], and [#15742] improve codec conversion, bandwidth extension, HF codec loading, and Lhotse codec data.
  • [#15524] adds semantic-codec training; [#15223] adds Frechet Codec Distance; [#15417] adds parallel scoring; [#15621] adds TTS comparison reports.
  • [#15348] and [#15427] improve multi-validation and deterministic multi-GPU eval.
  • [#15213], [#15212], [#15254], and [#15680] improve reference-free inference, dataset selection, inference config, and fp16 Whisper loading.

Audio

  • [#14442] adds a Conformer U-Net model for speech enhancement.
  • [#14626] fixes the Conformer U-Net implementation.
  • [#14749] adds a data-prediction objective for flow-matching speech-enhancement models.
  • [#14524] adds streaming mode to SpectrogramToAudio.
  • [#15388] fixes BNR 2.0 inference alignment with padded input signals.
  • [#15587] reduces duplication in the audio collection and applies small fixes.

DataLoader and Speech Data Tools

  • [#15303] adds Parquet/Arrow embedded-audio support via Lhotse.
  • [#15102] and [#15538] add AIS batch loading for ASR audio processing and LhotseSpeechToTextBpeDataset.
  • [#14891] supports input_cfg.yaml directly from AIS buckets.
  • [#15732] fixes non-tarred S3 audio in LazyNeMoIterator.
  • [#15200] adds on-the-fly temperature re-weighting for Lhotse datasets.
  • [#14558] adds randomized shard slicing for tarred data.
  • [#14586] adds in-manifest channel selection for multichannel recordings through Lhotse.
  • [#15316] and [#15410] add ShareGPT, indexed ShareGPT JSONL, and WebDataset data formats for SpeechLM2.
  • [#14294] adds ShareGPT data and a test loader.
  • [#15622] adds audio-codec Lhotse training format support; [#15742] updates and fixes audio-codec Lhotse loading.
  • [#14809] adds clipping, lowpass, and lossy-codec online augmentations; [#14808] adds augmented-audio saving.
  • [#15500] updates Speech Data Explorer with S3 read support, comparison mode, and security fixes; [#15137] updates the SDE tutorial.

Smaller Fixes

  • [#15437] adds .nemo torchaudio-preprocessor checkpoint migration.
  • [#15461] auto-detects use_bucketing and validates batch size.
  • [#15325] fixes TDT fusion scores; [#15322] moves transducer fusion models to the base class.
  • [#15245] speeds .nemo tar extraction/config processing.
  • [#15409] handles an exp-manager timer race.
  • [#15238] uses PurePosixPath.
  • [#15444] replaces assert with ValueError in EncDecMultiTaskModel.
  • [#15688] fixes validation audio logging; [#15699] adds UTMOSv2Calculator.verbose.
Source: README.md, updated 2026-08-06