Faster Whisper transcription with CTranslate2
Automatic Speech Recognition with Word-level Timestamps
Qwen3-TTS is an open-source series of TTS models
HunyuanVideo: A Systematic Framework For Large Video Generation Model
Fast State-of-the-Art Static Embeddings
Ultralytics YOLO
Achieving 3+ generation speedup on reasoning tasks
The most powerful local music generation model
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Native and Compact Structured Latents for 3D Generation
TokenSpeed is a speed-of-light LLM inference engine
Lets make video diffusion practical
A game theoretic approach to explain the output of ml models
100–200× Acceleration for Video Diffusion Models
Local long-term memory engine for AI apps with persistent storage
LightLLM is a Python-based LLM (Large Language Model) inference
High-Quality Voice Cloning TTS for 600+ Languages
A Web UI for easy subtitle using whisper model
An Open Source text-to-speech system built by inverting Whisper
Deep learning optimization library: makes distributed training easy
A text-to-speech, speech-to-text and speech-to-speech library
Ultra-Efficient LLMs on End Device
Fast and Universal 3D reconstruction model for versatile tasks
Free, high-quality text-to-speech API endpoint to replace OpenAI
Unified web UI for training and running open models locally