Document (PDF, Word, PPTX ...) extraction and parse API
OCR model for complex documents with layout-aware structured outputs
Generate audiobooks from EPUBs, PDFs and text with captions
Enhances Tesseract OCR output using LLMs (local or API)
A Repo For Document AI
Essential nodes that are weirdly missing from ComfyUI core
Misc; latest version of waifu2x; 2D video to stereo 3D video
PDF to Markdown with vision models
OCR software, free and offline
Faster Whisper transcription with CTranslate2
Stable Diffusion web UI
Python ETL framework for stream processing, real-time analytics, LLM
Open source healthcare AI
Persian NLP Toolkit
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Stable Diffusion web UI
Automatic subtitle synchronization tool
Translate the video from one language to another and embed dubbing
A full spaCy pipeline and models for scientific/biomedical documents
Visual Causal Flow
Comprehensive Gradio WebUI for audio processing
Cut videos with a text editor
Automated YouTube Shorts pipeline
Use Microsoft Edge's online text-to-speech service from Python
A TTS that fits in your CPU (and pocket)