Showing 64 open source projects for "vocal"

View related business solutions
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • Cut Data Warehouse Costs by 54% Icon
    Cut Data Warehouse Costs by 54%

    Easily migrate from Snowflake, Redshift, or Databricks with free tools.

    BigQuery delivers 54% lower TCO with exabyte scale and flexible pricing. Free migration tools handle the SQL translation automatically.
    Start Free
  • 1
    Ultimate Vocal Remover (UVR5)

    Ultimate Vocal Remover (UVR5)

    GUI for a Vocal Remover that uses Deep Neural Networks

    This application uses state-of-the-art source separation models to remove vocals from audio files. UVR's core developers trained all of the models provided in this package (except for the Demucs v3 and v4 4-stem models).
    Downloads: 900 This Week
    Last Update:
    See Project
  • 2
    YuE

    YuE

    Open source AI model for generating full songs from lyrics prompts

    YuE is an open source project that provides a foundation model designed for full-song music generation using artificial intelligence. It focuses on transforming text inputs such as lyrics and genre prompts into complete musical compositions that include both vocal and instrumental tracks. Unlike many shorter audio generators, the model is capable of producing songs that last several minutes while maintaining coherent musical structure and alignment with the provided lyrics. YuE introduces a family of models built on large language model architectures that process music generation as a sequence prediction task. ...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 3
    StemRoller

    StemRoller

    Isolate vocals, drums, bass, and other instrumental stems from songs

    StemRoller is the first free app that enables you to separate vocal and instrumental stems from any song with a single click! StemRoller uses Facebook's state-of-the-art Demucs algorithm for demixing songs and integrates search results from YouTube. Simply type the name/artist of any song into the search bar and click the Split button that appears in the results! You'll need to wait several minutes for splitting to complete.
    Downloads: 34 This Week
    Last Update:
    See Project
  • 4
    Transcoder

    Transcoder

    Hardware-accelerated video transcoding using Android MediaCodec APIs

    ...Unlike traditional speech translation systems that rely on multi-stage pipelines, Transcoder directly translates one speaker’s video into another language while preserving facial expressions, lip-sync, and vocal identity. Designed for real-time use and production-grade pipelines, Transcoder combines advanced deep learning models with GPU acceleration to deliver high-quality translations across languages. It’s built with researchers and developers in mind, offering tools for testing, evaluating, and deploying AI-driven media localization.
    Downloads: 3 This Week
    Last Update:
    See Project
  • Host LLMs in Production With On-Demand GPUs Icon
    Host LLMs in Production With On-Demand GPUs

    NVIDIA L4 GPUs. 5-second cold starts. Scale to zero when idle.

    Deploy your model, get an endpoint, pay only for compute time. No GPU provisioning or infrastructure management required.
    Start Free
  • 5
    Qwen2-Audio

    Qwen2-Audio

    Repo of Qwen2-Audio chat & pretrained large audio language model

    ...Code & examples provided with Hugging Face transformers, and usage via AutoProcessor, model classes etc. High performance on many standard benchmarks: ASR, speech-emotion recognition, vocal sound classification, speech translation etc.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 6
    ACE-Step 1.5

    ACE-Step 1.5

    The most powerful local music generation model

    ACE-Step 1.5 is an advanced open-source foundation model for AI-driven music generation that pushes beyond traditional limitations in speed, musical coherence, and controllability by innovating in architecture and training design. It integrates cutting-edge generative techniques—such as diffusion-based synthesis combined with compressed autoencoders and lightweight transformer elements—to produce high-quality full-length music tracks with rapid inference times, capable of generating a...
    Downloads: 71 This Week
    Last Update:
    See Project
  • 7
    GPT-SoVITS

    GPT-SoVITS

    1 min voice data can also be used to train a good TTS model

    GPT‑SoVITS is a state-of-the-art voice conversion and TTS system that enables zero‑shot and few‑shot synthesis based on a short vocal sample (e.g., 5 seconds). It supports cross‑lingual speech synthesis across English, Chinese, Japanese, Korean, Cantonese, and more. It's powered by VITS architecture enhanced for few‑sample adaptation and real‑time usability.
    Downloads: 13 This Week
    Last Update:
    See Project
  • 8
    Voice-Pro

    Voice-Pro

    Comprehensive Gradio WebUI for audio processing

    Voice-Pro is the best gradio WebUI for transcription, translation and text-to-speech. It can be easily installed with one click. Create a virtual environment using Miniconda, running completely separate from the Windows system (fully portable). Supports real-time transcription and translation, as well as batch mode.
    Downloads: 14 This Week
    Last Update:
    See Project
  • 9
    Step-Audio 2

    Step-Audio 2

    Multi-modal large language model designed for audio understanding

    ...It integrates a latent-space audio encoder, discrete acoustic tokens, and reinforcement-learning–based training (CoT + RL) to enhance its ability to capture and reproduce voice styles, intonations, and subtle vocal cues. Moreover, Step-Audio2 supports tool-calling and retrieval-augmented generation (RAG), allowing it to access external knowledge sources or audio/text databases, thus reducing hallucinations and improving coherence in complex dialogues.
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 10
    Step-Audio

    Step-Audio

    Open-source framework for intelligent speech interaction

    Step-Audio is a unified, open-source framework aimed at building intelligent speech systems that combine both comprehension and generation: it integrates large language models (LLMs) with speech input/output to handle not only semantic understanding but also rich vocal characteristics like tone, style, dialect, emotion, and prosody. The design moves beyond traditional separate-component pipelines (ASR → text model → TTS), instead offering a multimodal model that ingests speech or audio and produces speech accordingly, enabling natural dialogue, voice cloning, and expressive speech synthesis. Through its architecture, Step-Audio supports multilingual interaction, dialects, emotional tones (joy, sadness, etc.), and even more creative speech styles (like rap or singing), while allowing dynamic control over speech characteristics. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Fish Speech

    Fish Speech

    SOTA Open Source TTS

    ...Fish Speech emphasizes expressive and controllable voices: it supports a long list of emotion tags, tone markers, and special audio effect markers that can be embedded in the text to drive prosody and vocal style, from basic emotions to nuanced states like sarcastic, conciliative, or hysterical. The system is multilingual and cross-lingual, handling multiple languages in a single input without explicit phoneme markup, and is trained on large-scale datasets.
    Downloads: 15 This Week
    Last Update:
    See Project
  • 12
    Step-Audio-EditX

    Step-Audio-EditX

    LLM-based Reinforcement Learning audio edit model

    Step-Audio-EditX is an open-source, 3 billion-parameter audio model from StepFun AI designed to make expressive and precise editing of speech and audio as easy as text editing. Rather than treating audio editing as low-level waveform manipulation, this model converts speech into a sequence of discrete “audio tokens” (via a dual-codebook tokenizer) — combining a linguistic token stream and a semantic (prosody/emotion/style) token stream — thereby abstracting audio editing into high-level...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Admin365  AI Re-Edition

    Admin365 AI Re-Edition

    TechNews365 OS Admin AI intègre un Assistant Vocal IA 100% local !

    ...WinBoat : environnement Windows conteneurisé via Docker + RDP, isolation complète, compatible Cinnamon, installation et lancement validés. IA : Ollama + Phi‑3, Assistant Vocal IA (Whisper), Page Assist, Murmure . Admin & Réseau : GNS3, Wireshark, iperf3, phpIPAM, X2Go, AnyDesk, Docker + Compose, Portainer . Comment Mettre à jour le système? Ouvrez un terminal et exécutez : Support : 📩 contact@technews365.fr 🌐 https://technews365.fr
    Leader badge
    Downloads: 9 This Week
    Last Update:
    See Project
  • 14
    Free Karaoke File Maker

    Free Karaoke File Maker

    Free Karaoke File Maker

    You can hide the singer's voice from the music files that cannot hide the voice in the computer. By default, it will be saved with 2 audio tracks of singer + melody. If you want to save only the melody without the singer's voice, you have to select the No Vocal option. To save the output file, click Save Folder and choose the location you want to save (Default: Desktop). If you are sure of the above preparations, you can change the file you want to change by holding down the mouse and dragging it onto the Drag & Drop Input File. (No internet needed) You can also change it by clicking Select File.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    ResoRefine-AI

    ResoRefine-AI

    Free AI Audio Enhancer & Noise Reduction tool for Windows.

    ...•Remove Noise & Echo: Instantly clean up mic hiss, fan noise, and room echo using deep-learning suppression. •Studio Desk Control: Dial in "Broadcast Warmth" (Bass) and "Vocal Clarity" (Treble) for a rich, professional podcast sound. •Auto-Leveling: Built-in studio compression ensures perfectly balanced, clear audio for YouTube, Instagram, and TikTok.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 16
    byzorgan

    byzorgan

    Specialized sound synthesizer with Byzantine Church music scales

    This software integrates a small, specialized synthesizer and vocal processor. It can be used to learn Byzantine Church singing. You can play from the keyboard, mouse or touch screen. MIDI input is also available. Voice functions include: pitch highlighting, synthesizer control by voice, pitch correction and voice-to-ison conversion. On the screen there are labels with symbols of Byzantine notes.
    Downloads: 13 This Week
    Last Update:
    See Project
  • 17
    AI Local - Choose Models with LM Studio

    AI Local - Choose Models with LM Studio

    TN 365 AI Ready - LM Studio & Assistant Vocal Murmure !

    TN365 AI Edition est une version avancée de TN365, conçue pour offrir une expérience Linux moderne, rapide et entièrement orientée intelligence artificielle. Grâce à LM Studio préinstallé, les utilisateurs peuvent ajouter en un clic les modèles IA puissants tels que DeepSeek R1, Gemma, Llama 3.1, Qwen, Mistral, et bien d’autres. Pensée pour les créateurs, développeurs, Terminal DevOps inclus dans LM Studio, TN365 IA offre stabilité, élégance et puissance, tout en offrant une plateforme...
    Downloads: 12 This Week
    Last Update:
    See Project
  • 18
    TN 365 OS – Admin Edition 2026

    TN 365 OS – Admin Edition 2026

    TechNews365 OS Admin Edition – Tous les outils réseaux inclus

    TechNews365 OS Admin Edition – basée sur Linux Mint 22.3 Cinnamon Cette édition est conçue pour les administrateurs système, techniciens, dépanneurs et utilisateurs avancés. Elle fournit un environnement stable, rapide et complet, avec tous les outils nécessaires pour la gestion, le diagnostic, le réseau, la maintenance et la sécurité. ✔ Base : Linux Mint 22.3 Cinnamon ✔ Environnement : Cinnamon optimisé TechNews365 ✔ Architecture : x86_64 ✔ Outils Admin préinstallés ✔...
    Downloads: 10 This Week
    Last Update:
    See Project
  • 19
    vocal-separate

    vocal-separate

    An extremely simple tool for separating vocals and background music

    vocal-separate is a simple but effective audio processing application that isolates vocals and instrumental tracks from music and video files using stem-based source separation models, enabling tasks such as karaoke creation, remixing, and music analysis. Built as a localized web-based tool, it runs entirely on the user’s machine without requiring an internet connection, emphasizing privacy and convenience for creative work.
    Downloads: 5 This Week
    Last Update:
    See Project
  • 20
    TN 365 Maia 2026 et Mint-KDE

    TN 365 Maia 2026 et Mint-KDE

    Distribution TN 365 KDE moderne et stable !

    🇫🇷 TechNews365 OS Essentials – Édition KDE / (Base KDE Néon ou Mint-KDE) 2 ISO pour 2 environnements différents ! TechNews365 OS Essentials est une distribution Linux moderne, rapide et légère, basée sur KDE Néon et Mint-KDE . Elle offre une expérience simple, propre et optimisée pour le quotidien : KDE optimisé TN365 Thème Maia Transparent et icônes personnalisés Applications Essentielles incluse Radios, jeux légers, outils multimédia Calamares (installation...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 21
    RunningCoachFr

    RunningCoachFr

    Coach personnel francophone de courses fractionnées en musique!

    ⭐ Découvrez votre Coach personnel de courses fractionnées en musique ! 🏃🎵 ✨Créez facilement une session d'entraînement sur mesure avec votre propre playlist en MP3, guidée par un coach vocal francophone. Ce programme dynamique vous permet de concevoir des séances personnalisées pour optimiser vos performances en course fractionnée, tout en étant motivé par des musiques adaptées. ✨ ⚠️ Important - Windows SmartScreen Lorsque vous lancez ce logiciel, Windows peut afficher un message "Windows a protégé votre ordinateur". ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Demucs

    Demucs

    Code for the paper Hybrid Spectrogram and Waveform Source Separation

    Demucs (Deep Extractor for Music Sources) is a deep-learning framework for music source separation—extracting individual instrument or vocal tracks from a mixed audio file. The system is based on a U-Net-like convolutional architecture combined with recurrent and transformer elements to capture both short-term and long-term temporal structure. It processes raw waveforms directly rather than spectrograms, allowing for higher-quality reconstruction and fewer artifacts in separated tracks. ...
    Downloads: 97 This Week
    Last Update:
    See Project
  • 23
    Kalliope

    Kalliope

    Kalliope is a framework to create your own personal assistant

    ...Kalliope is a framework that will help you to create your own personal assistant. The concept is to create the brain of your assistant by attaching an input signal (vocal order, scheduled event, MQTT message, GPIO event, etc..) to one or multiple actions called neurons. You can create your own Kalliope bot, by simply choosing and composing the existing neurons without writing any code. But, if you need a particular module, you can write it by yourself, add it to your project, and propose it to the community. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    PseudonymizeSpeech

    PseudonymizeSpeech

    Praat script to pseudonymize speech.

    ...There is a trade-off between the level of pseudonymization and the (para-)linguistic features retained. The approach is to manipulate the spectro-temporal structure of the speech to simulate a different length and structure of the vocal tract, as well as a different pitch and speaking rate. The method is deterministic, and partially reversible. The extend of the changes is adjustable and gradual.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    KuStudio

    KuStudio

    Minimalistic and superstable OSC timeline/sequencer

    KuStudio is an open source OSC timeline sequencer, recorder and player, aimed to create timeline on an audiotrack. It can be used as core timeline module in interactive audiovisual and dance/vocal performances. For installation instructions see KuStudio-Guide.pdf, included in KuStudio archives. For quick support write to perevalovds@gmail.com KuStudio lets create, record and OSC tracks, synchronized with given audio track. Audiotrack can be WAV or AIFF file. KuStudio is inspired by famous Duration OSC editor, but has different philosophy: KuStudio stores all OSC tracks as discrete arrays, not curves, that allows to record and edit them freely. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • 3
  • Next