Control Any Computer Using LLMs
Document (PDF, Word, PPTX ...) extraction and parse API
High-performance inference server for text embeddings models API layer
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Python binding to the Apache Tika™ REST services
The most powerful and modular diffusion model GUI, api and backend
OCR software, free and offline
Stable Diffusion web UI
TTS with kokoro and onnx runtime
Generate audiobooks from e-books
A text-to-speech, speech-to-text and speech-to-speech library
Interface for OuteTTS models
A simple native web interface that uses ChatTTS to synthesize text
Python library and CLI tool to interface with Google Translate
Label Studio is a multi-type data labeling and annotation tool
A simple, high-quality voice conversion tool focused on ease of use
GUI for a Vocal Remover that uses Deep Neural Networks
Speech-AI-Forge is a project developed around TTS generation model
Awesome multilingual OCR toolkits based on PaddlePaddle
A Web UI for easy subtitle using whisper model
Self-host the powerful Chatterbox TTS model
Unified web UI for training and running open models locally
Generate audiobooks from e-books, voice cloning & 1107+ languages
Easily compute clip embeddings and build a clip retrieval system
LLM abstractions that aren't obstructions