Generate audiobooks from EPUBs, PDFs and text with captions
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Contexts Optical Compression
A nearly-live implementation of OpenAI's Whisper
OCR software, free and offline
Vim Win32 Installer
SoTA open-source TTS
A TTS that fits in your CPU (and pocket)
Official inference repo for FLUX.1 models
Offline Text To Speech synthesis for python
Tokenizer-Free TTS for Multilingual Speech Generation
State-of-the-art TTS model under 25MB
Code for running inference and finetuning with SAM 3 model
Robust Speech Recognition via Large-Scale Weak Supervision
Instant voice cloning by MIT and MyShell. Audio foundation model
Faster Whisper transcription with CTranslate2
Edit PDF files with Nano Banana
Open source annotation tool for machine learning practitioners
Mozc - a Japanese Input Method Editor designed for multi-platform
A Powerful Native Multimodal Model for Image Generation
Generate audiobooks from e-books
Industrial-level controllable zero-shot text-to-speech system
Ready-to-use OCR with 80+ supported languages
FastAPI framework, high performance, easy to learn, fast to code