modules free download

NVIDIA NeMo

Toolkit for conversational AI

...NeMo has separate collections for Automatic Speech Recognition (ASR), Natural Language Processing (NLP), and Text-to-Speech (TTS) models. Each collection consists of prebuilt modules that include everything needed to train on your data. Every module can easily be customized, extended, and composed to create new conversational AI model architectures. Conversational AI architectures are typically large and require a lot of data and compute for training. NeMo uses PyTorch Lightning for easy and performant multi-GPU/multi-node mixed-precision training. ...

Downloads: 2 This Week

Last Update: 2026-04-22

See Project

NVIDIA NeMo Framework

Scalable generative AI framework built for researchers and developers

NVIDIA NeMo is a scalable, cloud-native generative AI framework aimed at researchers and PyTorch developers working on large language models, multimodal models, and speech AI (ASR and TTS), with growing support for computer vision. It provides collections of domain-specific modules and reference implementations that make it easier to pre-train, fine-tune, and deploy very large models on multi-GPU and multi-node infrastructure. NeMo 2.0 introduces a Python-based configuration system, replacing YAML with more flexible, programmable configs that can be versioned and composed for different experiments. The framework builds on PyTorch Lightning–style modular abstractions, so training scripts are composed from reusable components for data loading, models, optimizers, and schedulers, which simplifies experimentation and adaptation. ...

Downloads: 0 This Week

Last Update: 2026-04-22

See Project

StyleTTS 2

Towards Human-Level Text-to-Speech through Style Diffusion

...StyleTTS2 supports both single-speaker and multi-speaker configurations, with the ability to sample or transfer styles from reference audio, making it powerful for expressive TTS and character voices. The repository includes training scripts, configuration files, and pre-trained auxiliary modules such as a text aligner, pitch extractor, and PL-BERT-based linguistic encoder.

Downloads: 4 This Week

Last Update: 2025-11-28

See Project

ekho

Chinese text-to-speech engine

...The code structure implies that Ekho may support hooking into audio input/output streams, perhaps for tasks like audio capture, playback, transformation, or simple voice-based operations. It might serve as a lightweight base or utility for building custom audio-related workflows, such as streaming, playback orchestration, or combining audio modules. Given the limited explicit features, Ekho would be best suited for developers or hobbyists who want a flexible foundation to add their own logic for TTS.

Downloads: 1 This Week

Last Update: 2025-11-28

See Project

Mocking Bird

Clone a voice in 5 seconds to generate arbitrary speech in real-time

...It builds on deep-learning based TTS / voice-cloning technology (in the lineage of projects such as Real-Time-Voice-Cloning), but extends it with support for Mandarin Chinese and multiple Chinese speech datasets — broadening its applicability beyond English. The codebase is implemented in Python (with PyTorch) and includes modules for encoder, synthesizer, vocoder, preprocessing, and inference, as well as demo scripts and a web-server interface for easier experimentation or deployment. MockingBird supports both using pretrained models and training your own synthesizer (with custom datasets), giving flexibility for voice-cloning or custom-voice synthesis depending on your needs.

1 Review

Downloads: 1 This Week

Last Update: 2023-03-23

See Project

Romanian Modular TTS

Modular Text-to-Speech system with a Matlab backbone. Your modules can be attached to this backbone via executable files (independent of the programming language used) respecting the XML interface requirements.

Downloads: 0 This Week

Last Update: 2013-04-11

See Project

FacialDAS

This project aims to distribute a facial animation system with speech, developed to brazilian portuguese case. This system is composed by many modules: movement extraction, facial animation and speech, through a text-to-speech system.

Downloads: 0 This Week

Last Update: 2015-09-22

See Project

Search Results for "modules"

Showing 7 open source projects for "modules"

NVIDIA NeMo

NVIDIA NeMo Framework

StyleTTS 2

ekho

Mocking Bird

Romanian Modular TTS

FacialDAS

Search Results for "modules"

Showing 7 open source projects for "modules"

NVIDIA NeMo

NVIDIA NeMo Framework

StyleTTS 2

ekho

Mocking Bird

Romanian Modular TTS

FacialDAS

Related Searches

Related Categories