Library for OCR-related tasks powered by Deep Learning
Comprehensive Gradio WebUI for audio processing
Automatic Speech Recognition with Word-level Timestamps
OCR software, free and offline
Generate audiobooks from EPUBs, PDFs and text with captions
A text-to-speech, speech-to-text and speech-to-speech library
Cut videos with a text editor
MiniMax H3 is a general-purpose, omni-modal generative system
A TTS that fits in your CPU (and pocket)
Generate audiobooks from e-books
SoTA open-source TTS
Official inference repo for FLUX.1 models
A simple native web interface that uses ChatTTS to synthesize text
MTEB: Massive Text Embedding Benchmark
Mozc - a Japanese Input Method Editor designed for multi-platform
A Powerful Native Multimodal Model for Image Generation
Powerful Android AI agent with tools, automation, and Linux shell
Controllable & emotion-expressive zero-shot TTS
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
A simple tool for reading in poorly redacted documents
Robust Speech Recognition via Large-Scale Weak Supervision
Open image model at the forefront of design
Offline Text To Speech synthesis for python
Faster Whisper transcription with CTranslate2
Snippet solution for Vim