ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Long-form streaming TTS system for multi-speaker dialogue generation
Interface for OuteTTS models
Automatic Speech Recognition with Word-level Timestamps
Open-source multi-speaker long-form text-to-speech model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
A generative speech model for daily dialogue
A Web UI for easy subtitle using whisper model
1B text generation model based on the HRM architecture
Buzz transcribes and translates audio offline
MOSS‑TTS Family open‑source speech and sound generation model
An Open Source implementation of Notebook LM with more flexibility
Play ChatGPT and other LLM with Xiaomi AI Speaker
Document (PDF, Word, PPTX ...) extraction and parse API
Oobabooga - The definitive Web UI for local AI, with powerful features
Official PyTorch Implementation
High-Quality Voice Cloning TTS for 600+ Languages
VoiceStudio is the open-source, fully-local ElevenLabs alternative
High-performance inference server for text embeddings models API layer
Hypernetworks that adapt LLMs for specific benchmark tasks
Provides line-oriented text file editing capabilities
Module for automatic summarization of text documents and HTML pages
Large Language Model Text Generation Inference
OpenVoiceOS Core, the FOSS Artificial Intelligence platform