Document (PDF, Word, PPTX ...) extraction and parse API
OCR model for complex documents with layout-aware structured outputs
Generate audiobooks from EPUBs, PDFs and text with captions
Enhances Tesseract OCR output using LLMs (local or API)
A Repo For Document AI
Essential nodes that are weirdly missing from ComfyUI core
Misc; latest version of waifu2x; 2D video to stereo 3D video
PDF to Markdown with vision models
OCR software, free and offline
Faster Whisper transcription with CTranslate2
Python ETL framework for stream processing, real-time analytics, LLM
Stable Diffusion web UI
Open source healthcare AI
Persian NLP Toolkit
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Stable Diffusion web UI
Automatic subtitle synchronization tool
A full spaCy pipeline and models for scientific/biomedical documents
Translate the video from one language to another and embed dubbing
Visual Causal Flow
Comprehensive Gradio WebUI for audio processing
Cut videos with a text editor
Automated YouTube Shorts pipeline
Use Microsoft Edge's online text-to-speech service from Python
A TTS that fits in your CPU (and pocket)