Tacotron is a heavily documented TensorFlow implementation of the end-to-end text-to-speech architecture introduced in the original Tacotron paper. It converts text into speech by learning acoustic representations and attention-based alignments from paired text and audio. The repository includes preprocessing, model modules, training, evaluation, and synthesis scripts. Example training setups use LJ Speech, Nick Offerman audiobook recordings, and the World English Bible dataset. Users can monitor loss and attention plots during training to evaluate alignment quality. Pretrained checkpoints and generated samples are provided as references for reproducing or studying the model's behavior.
Features
- End-to-end text-to-speech synthesis
- TensorFlow Tacotron implementation
- Audio and text preprocessing
- Training and evaluation scripts
- Attention alignment visualization
- Pretrained checkpoints and speech samples
Categories
AI ModelsLicense
Apache License V2.0Follow tacotron
Other Useful Business Software
$300 Free Credits to Build on Google Cloud
Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of tacotron!