Voice Recognition to Text Tool
EPUB to audiobook converter, optimized for Audiobookshelf
State-of-the-art TTS model under 25MB
Qwen3-TTS is an open-source series of TTS models
HunyuanVideo: A Systematic Framework For Large Video Generation Model
Official PyTorch Implementation
Python inference and LoRA trainer package for the LTX-2 audio–video
Framework for building realtime multimodal voice AI agents apps
AI-powered tool for generating, optimizing, and translating subtitles
A modular voice assistant application for experimenting
A lightweight text-to-speech model with zero-shot voice cloning
Towards Human-Sounding Speech
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Pythonic bindings for FFmpeg's libraries
The official Python Library for the Groq API
Python library and CLI tool to interface with Google Translate
Official repository for LTX-Video
LLM Large Model of Selling Anchor
Video editing with Python
A 0.1B Omni model trained from scratch
VMZ: Model Zoo for Video Modeling
High-resolution models for human tasks
Easy-to-use Speech Toolkit including Self-Supervised Learning model
Minimal scripts to run the emulator in a container for various systems
Get your documents ready for gen AI