VoiceStudio is an open-source, local-first platform for speech generation, voice cloning, transcription, dubbing, and audio production. It can clone voices from short reference recordings or design new voices from characteristics such as accent, pitch, age, and delivery style. Video dubbing combines transcription, translation, speaker preservation, speech synthesis, and final video export. It also supports multi-voice stories, audiobooks, dictation, vocal isolation, and speaker diarization. Multiple TTS and speech-recognition engines can run through CUDA, Apple Silicon, ROCm, or CPU hardware. Desktop, API, MCP, remote-worker, and batch-processing interfaces make it suitable for both interactive and automated workflows while keeping core processing local by default.

Features

  • Zero-shot voice cloning from short reference audio
  • Custom voice design with style and delivery controls
  • Automatic multilingual video dubbing
  • Audiobook and multi-voice story generation
  • Local transcription, diarization, and vocal isolation
  • Multiple TTS and ASR engines with GPU and CPU routing

Project Samples

Project Activity

See All Activity >

License

Affero GNU Public License

Follow VoiceStudio

VoiceStudio Web Site

Other Useful Business Software
Ship Agents Faster Icon
Ship Agents Faster

Transform your applications and workflows into powerful agentic systems at global scale.

Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of VoiceStudio!

Additional Project Details

Operating Systems

Linux, Mac, Windows

Programming Language

Python

Related Categories

Python Artificial Intelligence Software

Registered

2 days ago