Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Python library and CLI tool to interface with Google Translate
SQL-Driven RAG Engine
Translate the video from one language to another and embed dubbing
An Open Source text-to-speech system built by inverting Whisper
Multimodal-Driven Architecture for Customized Video Generation
Generate blog articles from video or audio
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Towards Human-Sounding Speech
Speech-AI-Forge is a project developed around TTS generation model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Handwritten Text Recognition (HTR) system implemented with TensorFlow
Reading book source
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
The simplest, fastest repository for training/finetuning models
High accuracy RAG for answering questions from scientific documents
A sound cloning tool with a web interface, using your voice
Open-source multi-speaker long-form text-to-speech model
Label Studio is a multi-type data labeling and annotation tool
Jupyter Notebooks as Markdown Documents, Julia, Python or R scripts
Snippet solution for Vim
Miso TTS is an 8 billion, highly emotive text-to-speech model
Powerful Android AI agent with tools, automation, and Linux shell
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
AI-powered tool for generating, optimizing, and translating subtitles