GUI for a Vocal Remover that uses Deep Neural Networks
Powerful Android AI agent with tools, automation, and Linux shell
Buzz transcribes and translates audio offline
The most powerful and modular diffusion model GUI, api and backend
Industry leading face manipulation platform
An enhanced tool for CodexApp, striving to make Codex better to use
Real time face swap and one-click video deepfake
Generate short videos with one click using AI LLM
State-of-the-art 2D and 3D Face Analysis Project
TTS with kokoro and onnx runtime
Advanced LLM-powered brute-force tool combining AI intelligence
MiniMax H3 is a general-purpose, omni-modal generative system
3D reconstruction software
Stable Diffusion web UI
VoiceStudio is the open-source, fully-local ElevenLabs alternative
OCRmyPDF adds an OCR text layer to scanned PDF files
Effortless data labeling with AI support from Segment Anything
Wan2.2: Open and Advanced Large-Scale Video Generative Model
A simple, high-quality voice conversion tool focused on ease of use
Faster Whisper transcription with CTranslate2
Awesome multilingual OCR toolkits based on PaddlePaddle
The agent that grows with you
The most powerful local music generation model
Robust Speech Recognition via Large-Scale Weak Supervision
Run Local LLMs on Any Device. Open-source