Automagically synchronize subtitles with video
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
EPUB to audiobook converter, optimized for Audiobookshelf
Automatically translates the text of a video based on a subtitle file
FAIR Sequence Modeling Toolkit 2
NLP Cloud serves high performance pre-trained or custom models for NER
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Persian NLP Toolkit
Bailing is a voice dialogue robot similar to GPT-4o
Scalable generative AI framework built for researchers and developers
Underthesea - Vietnamese NLP Toolkit
Multi-modal large language model designed for audio understanding
Official MiniMax Model Context Protocol (MCP) server
A 0.1B Omni model trained from scratch
Easy-to-use Speech Toolkit including Self-Supervised Learning model
A Web UI for easy subtitle using whisper model
Management of Yandex Station and other smart home devices
A speech-text foundation model for real time dialogue
Large Audio Language Model built for natural interactions
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Googles NotebookLM but local
Towards Studio-Grade Character Animation via In-Context Learning of 3D
Replace OpenAI GPT with another LLM in your app
Training data (data labeling, annotation, workflow) for all data types