Voice Recognition to Text Tool
Googles NotebookLM but local
Label Studio is a multi-type data labeling and annotation tool
Converts text to speech in realtime
EPUB to audiobook converter, optimized for Audiobookshelf
Turn words into chords
A high-quality rapid TTS voice cloning model
A TTS model capable of generating ultra-realistic dialogue
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Use Microsoft Edge's online text-to-speech service from Python
Open Source Speech Language Model
Qwen3-TTS is an open-source series of TTS models
SOTA discrete acoustic codec models with 40/75 tokens per second
Robust Speech Recognition via Large-Scale Weak Supervision
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
The official Python SDK for the ElevenLabs API
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Code and models for ICML 2024 paper, NExT-GPT
A Web UI for easy subtitle using whisper model
Sample code and notebooks for Generative AI on Google Cloud
Python library and CLI tool to interface with Google Translate
A simple native web interface that uses ChatTTS to synthesize text
Towards Human-Sounding Speech
An Open Source implementation of Notebook LM with more flexibility