Showing 150 open source projects for "speaker"

View related business solutions
  • Paessler - Monitor Your Whole Network in Minutes Icon
    Paessler - Monitor Your Whole Network in Minutes

    Auto-discovery finds your devices and deploys pre-configured sensors instantly. No project plan required, just visibility from day one.

    Waiting weeks for a monitoring rollout isn't an option when infrastructure doesn't stop running. PRTG's auto-discovery scans your network and suggests from over 200 pre-configured sensor types, so you're watching servers, applications and devices within minutes, not after a multi-week deployment. Enterprise-strength monitoring, without the enterprise complexity. Start your free trial today.
    Start Free 30-Day Trial
  • Earn up to 16% annual interest with Nexo. Icon
    Earn up to 16% annual interest with Nexo.

    Access competitive interest rates on your digital assets.

    Generate interest, borrow against your crypto, and trade a range of cryptocurrencies — all in one platform. Geographic restrictions, eligibility, and terms apply.
    Get started with Nexo.
  • 1
    tangent

    tangent

    Self-hosted Notes/To-Do App with voice transcription

    ...Your audio, transcripts, and notes never leave hardware you own — no cloud, no third party, no subscription. Features: Three capture modes — Brain Dump (verbatim voice memo), Meeting (speaker-aware transcript + meeting notes), Text Note (typed, no audio). Notebooks — a genuine ink surface: handwriting with pen styles, lasso selection, undo/redo, typed blocks and embedded recording cards on one scrolling page, and PDF export. Handwriting search — search all your notebooks in a flash Ask My Notes — questions grounded in your recordings, summaries, notebooks, and to-dos, with actionable items produced from 2 second voice note
    Downloads: 7 This Week
    Last Update:
    See Project
  • 2
    IMYplay

    IMYplay

    Plays iMelody (IMY) files using many sound systems

    ...PortAudiov19 (http://www.portaudio.com), 7. PulseAudio (http://www.pulseaudio.org), 8. JACK1/JACK2 (http://jackaudio.org), 9. GStreamer (http://gstreamer.freedesktop.org), 10. PC-Speaker (at least Linux and DOS). It can also: - convert IMY ringtones to MIDI files, - convert IMY ringtones to WAV files, - write raw samples to an output file, - execute an external program on each note. See the project homepage https://imyplay.sourceforge.io and the project Wiki in the menu above.
    Leader badge
    Downloads: 10 This Week
    Last Update:
    See Project
  • 3
    Conversations

    Conversations

    App in java for chatting to a generative A.I. (involving tts and stt)

    ... * The AI ​​responds and the server returns that response in real time, and the sentences converted to audio (textToSpeech), and the application broadcasts them through the speaker. The application is prepared so that only one user occupies the server's resources, so if the server is busy, in theory it will not let you connect. There is a demo video that shows how it works: https://frojasg1.com:8443/resource_counter/resourceCounter?operation=countAndForward&url=https%3A%2F%2Ffrojasg1.com%2Fdemos%2Faplicaciones%2Fchat%2F20240815.Demo.Chat.mp4%3Forigin%3Dsourceforge&origin=web More about it at this web site: https://www.frojasg1.com:8443/downloads_web/web/html/conversaciones.html?...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 4
    MiroTalk SFU

    MiroTalk SFU

    🏆 MiroTalk SFU - WebRTC video call, chat, screen sharing & more

    Open Source Project, Free and Self-Hosted. A good alternative of Zoom, Google Meet, Microsoft Teams, GoToMeeting... More details here: https://github.com/miroslavpejic85/mirotalksfu
    Downloads: 1 This Week
    Last Update:
    See Project
  • Host LLMs in Production With On-Demand GPUs Icon
    Host LLMs in Production With On-Demand GPUs

    NVIDIA L4 GPUs. 5-second cold starts. Scale to zero when idle.

    Deploy your model, get an endpoint, pay only for compute time. No GPU provisioning or infrastructure management required.
    Start Free
  • 5
    Bluetooth Internet Radio

    Bluetooth Internet Radio

    Auto-setup bluetooth internet radio on Ubuntu / Raspberry Pi / Armbian

    Fully automated setup of bluetooth, internet radio player (PyRadio), local music files player (mplayer), on a headless RaspberryPi, connecting to a stereo system's bluetooth receiver (bash script, chmod +x it to run). To install automatically, copy => paste => run the commands below in a terminal program (using the 'Terminal' app in the system menu, or over remote SSH), while logged in AS THE USER THAT WILL RUN THE APP (user must have sudo privileges): wget --no-cache -O...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    MetaVoice-1B

    MetaVoice-1B

    Foundational model for human-like, expressive TTS

    MetaVoice — in the form of its source repository “metavoice-src” — is a large-scale text-to-speech (TTS) model. Specifically, the base model (MetaVoice-1B) uses around 1.2 billion parameters and has been trained on a massive dataset — reportedly around 100,000 hours of speech data. The goal is to provide human-like, expressive, and flexible TTS: able to generate natural-sounding speech that can handle diverse inputs and likely generalize over voice styles, intonation, prosody, and perhaps...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    ChatTTS_colab

    ChatTTS_colab

    One-click deployment (including offline integration package)

    ...It has first-class support for long-form audio generation, making it suitable for audiobooks, podcasts, or long narration tasks. The project also implements multi-speaker or role-based reading, letting users assign different voices to different characters in a script and even use a large language model to generate that script in one step.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    footswitch2

    footswitch2

    Audio Transcription software for Linux (Vlc) with a foot pedal

    Footswitch 2 is a media player for transcribers on Linux. Written in python and using the python bindings for VLC it allows a transcriber to control the audio or video with a USB footpedal, and includes a set of macros that integrate into LibreOffice. This allows the transcriber to control the media player from within Libreoffice as well, making it useful for those who do not yet own a footpedal/footswitch. Control of the media player from LibreOffice can be via Hotkeys or an integrated...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 9
    MARS5

    MARS5

    MARS5 speech model (TTS) from CAMB.AI

    ...The model is built to handle prosodically challenging content such as sports commentary, anime dialogue, and other high-energy or highly varied speech patterns with realistic rhythm and intonation. To control speaker identity, MARS5 uses a short reference audio clip, typically between 2 and 12 seconds, from which it learns the voice characteristics. It supports two main inference modes: shallow clone, which is faster and only needs the reference audio, and deep clone, which additionally uses the transcript of the reference audio to increase similarity and naturalness at the cost of more computation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • 10
    SoniTranslate

    SoniTranslate

    Synchronized Translation for Videos

    SoniTranslate is a video translation and dubbing system that produces synchronized target-language audio tracks for existing video content. It provides a web UI built with Gradio, allowing users to upload a video, choose source and target languages, and then run a pipeline that handles transcription, translation and re-synthesis of speech. Under the hood, it uses advanced speech and diarization models to separate speakers, align audio with timecodes and respect subtitle timing, which lets...
    Downloads: 38 This Week
    Last Update:
    See Project
  • 11
    StyleTTS 2

    StyleTTS 2

    Towards Human-Level Text-to-Speech through Style Diffusion

    ...The architecture uses a two-stage training process and leverages an auxiliary speech language model to guide generation toward more natural and coherent utterances. StyleTTS2 supports both single-speaker and multi-speaker configurations, with the ability to sample or transfer styles from reference audio, making it powerful for expressive TTS and character voices. The repository includes training scripts, configuration files, and pre-trained auxiliary modules such as a text aligner, pitch extractor, and PL-BERT-based linguistic encoder.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    GraphQL SPQR

    GraphQL SPQR

    Build a GraphQL service in seconds

    GraphQL SPQR (GraphQL Schema Publisher & Query Resolver, pronounced like speaker) is a simple-to-use library for rapid development of GraphQL APIs in Java. GraphQL SPQR aims to make it dead simple to add a GraphQL API to any Java project. It works by dynamically generating a GraphQL schema from Java code. When developing GraphQL-enabled applications it is common to define the schema first and hook up the business logic later.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    VALL-E X

    VALL-E X

    Open source implementation of Microsoft's VALL-E X zero-shot TTS model

    VALL-E-X is an open-source implementation of Microsoft’s VALL-E X zero-shot text-to-speech model, focused on multilingual, cross-lingual voice cloning. It is capable of synthesizing speech in English, Chinese, and Japanese from text while mimicking the voice characteristics of a speaker given only a short 3–10 second prompt. The model attempts to match not just timbre, but also tone, pitch, emotion, and prosody of the reference audio, resulting in highly personalized output. VALL-E-X supports zero-shot cross-lingual synthesis, meaning a monolingual speaker’s voice can be used to speak other languages without additional training. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    wukong-robot

    wukong-robot

    Chinese voice dialogue robot/smart speaker project

    wukong-robot is a Chinese voice assistant / smart speaker project built to let makers and hackers design highly customizable voice-controlled devices. It combines wake-word detection, automatic speech recognition, natural language understanding, and text-to-speech into a single framework aimed at the Chinese-speaking ecosystem. The project is positioned as a simple, flexible, and elegant platform that can run on devices like Raspberry Pi and other Linux-based boards, making it suitable for DIY smart speakers and home-automation hubs. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 15
    lora-svc

    lora-svc

    Singing voice change based on whisper, lora for singing voice clone

    singing voice change based on whisper, and lora for singing voice clone. You will feel the beauty of the code from this project. Uni-SVC main branch is for singing voice clone based on whisper with speaker encoder and speaker adapter. Uni-SVC main target is to develop lora for SVC. With lora, maybe clone a singer just need 10 stence after 10 minutes train. Each singer is a plug-in of the base model.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    VALL-E

    VALL-E

    PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)

    ...VALL-E emerges in-context learning capabilities and can be used to synthesize high-quality personalized speech with only a 3-second enrolled recording of an unseen speaker as an acoustic prompt. Experiment results show that VALL-E significantly outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity. In addition, we find VALL-E could preserve the speaker's emotion and acoustic environment of the acoustic prompt in synthesis.
    Downloads: 4 This Week
    Last Update:
    See Project
  • 17

    cwdaemon

    morse daemon for the serial or parallel port

    .... ------ cwdaemon is a small daemon which uses the PC parallel or serial port and a simple transistor switch to output morse code to a transmitter from a text message sent to it via the udp internet protocol. The program uses the soundcard or PC speaker to generate a sidetone. Current version of cwdaemon is 0.12.0, it has been released in May 2023. Source code for cwdaemon is available on github: https://github.com/acerion/cwdaemon
    Downloads: 2 This Week
    Last Update:
    See Project
  • 18
    footswitch3

    footswitch3

    Audio Transcription software for Linux (Gstreamer) with a foot pedal

    Footswitch 3 is a media player for transcribers on Linux. Written in python using the python bindings for Gstreamer it allows a transcriber to control the audio or video with a foot pedal, and includes a set of macros that integrate into LibreOffice. This allows the transcriber to control the media player from within Libreoffice as well, making it useful for those who do not yet own a foot pedal/foot switch. Control of the media player from LibreOffice can be via Hotkeys or an integrated...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 19
    Loudspeaker Design Calculations Toolkit

    Loudspeaker Design Calculations Toolkit

    LDCT is a cross-platform tool for common speaker design calculations.

    LDCT is an open source, cross-platform application containing many calculations commonly used in loudspeaker design.
    Downloads: 9 This Week
    Last Update:
    See Project
  • 20

    ShiftRC

    Simple cheap DIY RC receiver/handler.

    ...this project will host multiple version, some very simple ones, and some more complex. a similar thing can be said for the controllers, I will soon upload the schematics and designs for a steampunk style controller. but for a more simple understanding a simple remote might be a transmitter(led, speaker, or antenna for example) a button, and a capacitor to stabilize the button because every pulse you send changes the state. however high precission control is possible. Use the project and schematics however you want it as long as it does NOT restrict others.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    SVoice (Speech Voice Separation)

    SVoice (Speech Voice Separation)

    We provide a PyTorch implementation of the paper Voice Separation

    SVoice is a PyTorch-based implementation of Facebook Research’s study on speaker voice separation as described in the paper “Voice Separation with an Unknown Number of Multiple Speakers.” This project presents a deep learning framework capable of separating mixed audio sequences where several people speak simultaneously, without prior knowledge of how many speakers are present. The model employs gated neural networks with recurrent processing blocks that disentangle voices over multiple computational steps, while maintaining speaker consistency across output channels. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    reveal-jekyll

    reveal-jekyll

    Online presentation for GitHub Pages and Jekyll in Markdown

    Transforms Markdown files into presentation slides using reveal.js and Jekyll. The theme is based on Solarized Colors (by Ethan Schoonover) containing a light and a dark theme. reveal-jekyll is ready for GitLab Pages as well as GitHub Pages. Jekyll is a simple, blog-aware, static site generator perfect for personal, project, or organization sites. Think of it like a file-based CMS, without all the complexity. Jekyll takes your content, renders Markdown and Liquid templates, and spits out a...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 23
    VITS

    VITS

    Conditional Variational Autoencoder with Adversarial Learning

    ...This architecture enables parallel generation (fast inference) while achieving speech quality that rivals or surpasses many two-stage systems. The repository provides training and inference pipelines for common datasets such as LJ Speech (single-speaker) and VCTK (multi-speaker), including filelists, configs, and preprocessing scripts. It also includes monotonic alignment search code and g2p preprocessing, which are crucial components for aligning text and speech in an end-to-end setup.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 24
    FUSUMA

    FUSUMA

    Fusuma makes slides with Markdown easily

    ...Development Mode, for running with HMR, just coding Markdown and sometimes CSS. Build Mode, to render to HTML and optimizing js,css,md. Generating an image of slides as og:image and checking a11y automatically. Presentation Mode, including speaker note, timer, and features for recording your page actions and voice. Deploy Mode, for deploying to GitHub pages. And PDF Mode, for exporting slides as PDF.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    TTS

    TTS

    Deep learning for text to speech

    ...Notebooks for extensive model benchmarking. Modular (but not too much) code base enabling easy testing for new ideas. Text2Spec models (Tacotron, Tacotron2, Glow-TTS, SpeedySpeech). Speaker Encoder to compute speaker embeddings efficiently. Vocoder models (MelGAN, Multiband-MelGAN, GAN-TTS, ParallelWaveGAN, WaveGrad, WaveRNN). If you are only interested in synthesizing speech with the released TTS models, installing from PyPI is the easiest option.
    Downloads: 0 This Week
    Last Update:
    See Project