Showing 150 open source projects for "speaker"

View related business solutions
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    MOSS-TTS Family

    MOSS-TTS Family

    MOSS‑TTS Family open‑source speech and sound generation model

    MOSS-TTS is an open-source speech and sound generation model family built for high-fidelity, expressive, and production-oriented audio workflows. It covers long-form speech, voice cloning, multi-speaker dialogue, voice design, environmental sound effects, and real-time streaming TTS. The project is designed for complex real-world use cases where a single speech model may not be enough. Its flagship model focuses on stable long speech generation, multilingual and code-switched synthesis, pronunciation control, and zero-shot voice cloning. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 2
    Note67

    Note67

    A private, local meeting notes assistant

    ...Built with a cross-platform architecture using Rust (via Tauri) for backend logic and a TypeScript/React frontend, it prioritizes privacy by performing audio transcription locally with Whisper models and generating summaries with locally-hosted AI, eliminating the need to send sensitive meeting content to external servers. Users can record meetings directly from their microphone, view live transcriptions, filter by speaker, and export structured summaries, making it useful for professionals who need searchable, organized records of discussions. It also features thoughtful signal processing such as voice activity detection and echo deduplication to improve transcription accuracy, and provides standard note-taking features.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    AutoSubs

    AutoSubs

    Instantly generate AI-powered subtitles on your device

    AutoSubs is an open-source, AI-powered subtitle generation tool that enables users to automatically transcribe audio and video content into accurate, editable subtitles directly on their device. It supports both standalone usage and integration with professional video editing software such as DaVinci Resolve, allowing creators to generate and edit subtitles within their existing workflows. The tool leverages speech-to-text models, including OpenAI Whisper, to produce high-quality...
    Downloads: 59 This Week
    Last Update:
    See Project
  • 4
    ChatTTS

    ChatTTS

    A generative speech model for daily dialogue

    ChatTTS is an open-source conversational text-to-speech model optimized for dialogue, developed by 2Noise. Trained on 100,000+ hours of English and Chinese conversation data, it excels at generating expressive prosody—pauses, interjections, laughter—for more natural-sounding speech synthesis in assistant and chatbot applications.
    Downloads: 5 This Week
    Last Update:
    See Project
  • Go from Code to Production URL in Seconds Icon
    Go from Code to Production URL in Seconds

    Cloud Run deploys apps in any language instantly. Scales to zero. Pay only when code runs.

    Skip the Kubernetes configs. Cloud Run handles HTTPS, scaling, and infrastructure automatically. Two million requests free per month.
    Start Free
  • 5
    pyVideoTrans

    pyVideoTrans

    Translate the video from one language to another and embed dubbing

    pyVideoTrans is an ambitious open-source multimedia processing project that assembles speech recognition, subtitle generation, AI translation, voice synthesis, and video assembly into a unified pipeline for converting videos from one language to another with embedded dubbing and captions. At its core it runs speech-to-text models to transcribe audio tracks, translates the resulting text into a target language using local or cloud-based translation engines, synthesizes new speech to match the...
    Downloads: 28 This Week
    Last Update:
    See Project
  • 6
    Local-NotebookLM

    Local-NotebookLM

    Googles NotebookLM but local

    Local-NotebookLM is a local AI tool for turning PDF documents into generated audio content. It works like a self-hosted alternative to NotebookLM-style document-to-audio workflows. The system extracts and processes PDF text, sends the content through an LLM, and converts the result into speech with configurable voices. Users can generate podcasts, summaries, interviews, lectures, debates, tutorials, news reports, executive briefs, and other formats. It supports multiple LLM providers,...
    Downloads: 18 This Week
    Last Update:
    See Project
  • 7
    Bento

    Bento

    Bento, the office suite that fits in a file

    ...It opens in a modern browser without an account or installer and rewrites its own embedded data when saved. The document model remains readable as plain JSON, making files portable, inspectable, and directly editable by AI agents. Bento includes morph transitions, speaker view, comments, motion paths, PDF export, and built-in interactive charts. Optional collaboration uses end-to-end encryption and a CRDT so offline edits can merge while the relay stores only ciphertext. An offline mode blocks updates and collaboration, while signed self-updates create a new file and preserve the previous copy for rollback.
    Downloads: 24 This Week
    Last Update:
    See Project
  • 8
    Music Assistant

    Music Assistant

    Music Assistant is a free, opensource Media library manager

    Music Assistant Server is the core backend for Music Assistant, a free and open-source music library manager for local and online music sources. It connects streaming services, local files, metadata providers, and many speaker ecosystems into one centralized music system. The server is designed to run on an always-on device such as a Raspberry Pi, NAS, Intel NUC, or similar home server. It can work as a standalone product, but it is especially tailored for Home Assistant users who want automation, voice control, and smart-home playback workflows. Music Assistant supports features such as library matching, metadata enrichment, gapless playback, crossfade, volume normalization, synchronized playback, announcements, and queue transfers. ...
    Downloads: 21 This Week
    Last Update:
    See Project
  • 9
    Step-Audio 2

    Step-Audio 2

    Multi-modal large language model designed for audio understanding

    Step-Audio2 is an advanced, end-to-end multimodal large language model designed for high-fidelity audio understanding and natural speech conversation: unlike many pipelines that separate speech recognition, processing, and synthesis, Step-Audio2 processes raw audio, reasons about semantic and paralinguistic content (like emotion, speaker characteristics, non-verbal cues), and can generate contextually appropriate responses — including potentially generating or transforming audio output. It integrates a latent-space audio encoder, discrete acoustic tokens, and reinforcement-learning–based training (CoT + RL) to enhance its ability to capture and reproduce voice styles, intonations, and subtle vocal cues. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Paessler: Easy to Use With Enterprise Power. Free Trial Icon
    Paessler: Easy to Use With Enterprise Power. Free Trial

    A low-code dashboard makes monitoring intuitive for any admin, while scripting and custom sensors give experts full control.

    You shouldn't have to choose between a monitoring tool that's easy to use and one that's powerful enough for a complex environment. PRTG's low-code interface lets any admin build dashboards, set alerts and monitor devices without scripting, while custom sensors and full API access are there when your team needs deeper control. One platform, no compromise. Download a free 30-day trial now.
    Get Free Download
  • 10
    OmniVoice

    OmniVoice

    High-Quality Voice Cloning TTS for 600+ Languages

    The OmniVoice project is a cutting-edge multilingual text-to-speech system designed to generate high-quality speech across more than 600 languages. Built on a diffusion language model-style architecture, it combines scalability with strong performance, enabling both natural-sounding voice synthesis and efficient inference speeds. One of its most notable capabilities is zero-shot voice cloning, allowing users to replicate a speaker’s voice using only a short reference audio clip. In addition,...
    Downloads: 19 This Week
    Last Update:
    See Project
  • 11
    jekyll-theme-conference

    jekyll-theme-conference

    Jekyll template for a conference website containing program

    This is a responsive Jekyll theme based on Bootstrap 4 for conferences. All components such as talks, speakers or rooms are represented as collection of files. The schedule is given is defined via a simple structure stored in a YAML file. There is no need for databases and once generated the website consists only of static files. A script and workflows are available for easy import, e.g. of frab-compatible schedules. The design is easily customizable and is adapted for mobile uses and...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    OpenVoice

    OpenVoice

    Instant voice cloning by MIT and MyShell. Audio foundation model

    ...It is designed not only to match the timbre of the reference voice, but also to give granular control over style parameters such as emotion, accent, rhythm, pauses, and intonation. The model supports cross-lingual and even zero-shot cross-lingual voice cloning, so a speaker recorded in one language can be made to speak naturally in others. Architecturally, OpenVoice separates “tone color” cloning from style control, which makes it easier to keep a consistent identity while flexibly changing prosody or language. The project provides open-weight models, inference code, and examples, making it suitable both for research and for building production voice experiences. ...
    Downloads: 14 This Week
    Last Update:
    See Project
  • 13
    SpatGRIS4

    SpatGRIS4

    SpatGRIS4

    ... - Various adjustments to settings of the trajectories based on audio signal analysis. SpeakerView 1.0.3 - Ability to display the Triplets for Dome and Hybrid modes in binaural listening, in addition to accessing the speaker levels (View>Show Speaker Levels)
    Downloads: 64 This Week
    Last Update:
    See Project
  • 14
    MiMo-V2.5-ASR

    MiMo-V2.5-ASR

    Robust Speech Recognition Across Languages, Dialects

    MiMo-V2.5-ASR is an advanced automatic speech recognition system developed as part of Xiaomi’s MiMo AI ecosystem. It is designed to handle complex acoustic environments, including noisy conditions and diverse speaker variations. The model supports multiple languages and dialects, enabling robust transcription across global use cases. It leverages modern deep learning architectures to improve accuracy and adaptability in real-world scenarios. The system is built to integrate with broader AI pipelines, including voice assistants and multimodal systems. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Fulcro

    Fulcro

    A library for development of single-page full-stack web applications

    Fulcro is a batteries-included, full-stack library for building data-driven web applications in Clojure and ClojureScript. It integrates seamlessly with React, supports web, native, and desktop (Electron) targets, and enables strong local reasoning while facilitating rapid development and scalable production-level UIs. The rewrite of Fulcro Inspect is available via the releases page of Fulcro Inspect. And there are preliminary instructions for using it with the latest Fulcro. The alpha...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    PowerPoint-ist

    PowerPoint-ist

    Web presentation editor replicating many PowerPoint features online

    PPTist is a web-based presentation editing application designed to replicate many of the commonly used features found in traditional slide presentation software. It allows users to create, edit, and present slide decks directly within a web browser while maintaining a desktop-like editing experience. PPTist is built with Vue 3 and TypeScript and focuses on providing a highly interactive slide editing environment with extensive customization and extension potential. PPTist supports a wide...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 17
    Spark TTS

    Spark TTS

    Spark-TTS Inference Code

    ...The project supports zero-shot voice cloning, meaning it can imitate a new speaker’s voice without dedicated training for that specific voice, and works across languages, including English and Chinese, even in cross-lingual code-switching scenarios. Spark-TTS allows users to control speech characteristics like gender, pitch, and speaking rate to customize synthesized output and support virtual speaker creation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    ESPnet

    ESPnet

    End-to-end speech processing toolkit

    ESPnet is a comprehensive end-to-end speech processing toolkit covering a wide spectrum of tasks, including automatic speech recognition (ASR), text-to-speech (TTS), speech translation (ST), speech enhancement, speaker diarization, and spoken language understanding. It uses PyTorch as its deep learning engine and adopts a Kaldi-style data processing pipeline for features, data formats, and experimental recipes. This combination allows researchers to leverage modern neural architectures while still benefiting from the robust data preparation practices developed in the speech community. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Audio Share

    Audio Share

    Audio Share can share Windows/Linux computer's audio to Android phone

    Audio Share can share Windows/Linux computer's audio to Android phone over network, so your phone becomes the speaker of computer. (You needn't buy a new speaker😄.) https://github.com/mkckr0/audio-share
    Leader badge
    Downloads: 383 This Week
    Last Update:
    See Project
  • 20

    EasyBluetoothAudio

    Turn your Windows PC into a high-fidelity Bluetooth Speaker.

    Downloads: 3 This Week
    Last Update:
    See Project
  • 21
    Reflex

    Reflex

    Audio measurement sfotware

    Qt-based software for audio measurement (speaker parameters, frequency and phase response, etc).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    JSpeech

    JSpeech

    Java library designed to integrate Speech-to-Text

    jSpeech is a Java library designed to integrate Speech-to-Text (STT) capabilities, command control, and diarization (speaker identification) into applications in a simple, modular, and decoupled way.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 23
    SpatGRIS

    SpatGRIS

    The SpatGRIS is an external multichannel sound spatializer.

    ...https://sourceforge.net/projects/spatgris4/ October 2024: Version 3.3.7. This is the final version of SpatGRIS3. • SpatGRIS is now Universal (Intel and Apple Silicon) New: • A general MUTE button • BlackHole 0.6.0 • Speaker Setup 3D view is now a separate application: SpeakerView • French Manual • SpatGRIS PLAYER: a function that allows using SpatGRIS3 as a standalone software to play any piece encoded by it to any speaker setup. • SpatGRIS can now be used in "HYBRID" mode. When activated, each audio source can be set to either "DOME" or "CUBE"...
    Leader badge
    Downloads: 10 This Week
    Last Update:
    See Project
  • 24
    MUN Timer

    MUN Timer

    A Customizable GUI timer designed for MUN.

    MUN Timer Pro: Ultimate Model United Nations App Streamline your Model UN sessions with MUN Timer Pro, a flexible timer app for chairs and delegates. Manage agendas and timelines effortlessly in an intuitive interface. Perfect for MUN conferences, debates, and simulations. Boost efficiency—download now!
    Downloads: 2 This Week
    Last Update:
    See Project
  • 25
    Accessible-Coconut

    Accessible-Coconut

    A GNU/Linux operating system accessible for visually impaired.

    Accessible-Coconut(AC) is a community driven GNU/Linux operating system which is completely accessible for persons with visual impairment. AC is derived from Ubuntu-MATE. Yes the goal is to make a free and open-source eyes free desktop environment. Forum : https://groups.google.com/forum/#!forum/accessible-coconut Telegram forum : https://telegram.me/accessible_coconut Project home : https://zendalona.com/
    Leader badge
    Downloads: 33 This Week
    Last Update:
    See Project