1B text generation model based on the HRM architecture
Document (PDF, Word, PPTX ...) extraction and parse API
Oobabooga - The definitive Web UI for local AI, with powerful features
High-performance inference server for text embeddings models API layer
Hypernetworks that adapt LLMs for specific benchmark tasks
Provides line-oriented text file editing capabilities
Module for automatic summarization of text documents and HTML pages
Large Language Model Text Generation Inference
TTS with kokoro and onnx runtime
AI tool that removes hardcoded subtitles and text from videos locally
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
OCRmyPDF adds an OCR text layer to scanned PDF files
Awesome multilingual OCR toolkits based on PaddlePaddle
Comprehensive Gradio WebUI for audio processing
Automatic Speech Recognition with Word-level Timestamps
High-Quality Voice Cloning TTS for 600+ Languages
Focus on prompting and generating
SOTA Open Source TTS
Python tool for converting files and office documents to Markdown
Qwen3-TTS is an open-source series of TTS models
A generative speech model for daily dialogue
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Generate audiobooks from EPUBs, PDFs and text with captions
A nearly-live implementation of OpenAI's Whisper
Wan2.1: Open and Advanced Large-Scale Video Generative Model