VoiceStudio is an open-source, local-first platform for speech generation, voice cloning, transcription, dubbing, and audio production. It can clone voices from short reference recordings or design new voices from characteristics such as accent, pitch, age, and delivery style. Video dubbing combines transcription, translation, speaker preservation, speech synthesis, and final video export. It also supports multi-voice stories, audiobooks, dictation, vocal isolation, and speaker diarization. Multiple TTS and speech-recognition engines can run through CUDA, Apple Silicon, ROCm, or CPU hardware. Desktop, API, MCP, remote-worker, and batch-processing interfaces make it suitable for both interactive and automated workflows while keeping core processing local by default.
Features
- Zero-shot voice cloning from short reference audio
- Custom voice design with style and delivery controls
- Automatic multilingual video dubbing
- Audiobook and multi-voice story generation
- Local transcription, diarization, and vocal isolation
- Multiple TTS and ASR engines with GPU and CPU routing