OCRmyPDF adds an OCR text layer to scanned PDF files
Python library and CLI tool to interface with Google Translate
Use Microsoft Edge's online text-to-speech service from Python
A generative speech model for daily dialogue
MiniMax H3 is a general-purpose, omni-modal generative system
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
A text-to-speech, speech-to-text and speech-to-speech library
High-Quality Voice Cloning TTS for 600+ Languages
Awesome multilingual OCR toolkits based on PaddlePaddle
Official inference repo for FLUX.1 models
Comprehensive Gradio WebUI for audio processing
Extensions for Python Markdown
A robust, efficient, low-latency speech-to-text library
Generate audiobooks from EPUBs, PDFs and text with captions
Comprehensive Markdown plugin built for Django
State-of-the-art TTS model under 25MB
Python & command-line tool to gather text on the Web
Cut videos with a text editor
A TTS that fits in your CPU (and pocket)
Hyperextensible Vim-based text editor
ASCII art library for Python
Jupyter Notebooks as Markdown Documents, Julia, Python or R scripts
A simple native web interface that uses ChatTTS to synthesize text
Automatic Speech Recognition with Word-level Timestamps
Offline inference engine for art, real-time voice conversations