OCR software, free and offline
Faster Whisper transcription with CTranslate2
Generate audiobooks from EPUBs, PDFs and text with captions
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Comprehensive Gradio WebUI for audio processing
Robust Speech Recognition via Large-Scale Weak Supervision
Use Microsoft Edge's online text-to-speech service from Python
A TTS that fits in your CPU (and pocket)
Open source healthcare AI
OpenAI gpt-image-2 API
Unlimited, private and free Speech-To-Text program
1 min voice data can also be used to train a good TTS model
Stable Diffusion web UI
AI tool for automatic batch short video creation and editing
A modular voice assistant application for experimenting
Contexts Optical Compression
AsrTools: Smart Voice-to-Text Tool
Visual Causal Flow
EPUB to audiobook converter, optimized for Audiobookshelf
Python library and CLI tool to interface with Google Translate
95% token savings. 155x faster queries. 16 languages
OCR model for complex documents with layout-aware structured outputs
Free, high-quality text-to-speech API endpoint to replace OpenAI
Implementation of Imagen, Google's Text-to-Image Neural Network
Automated translation solution for visual novels