Document (PDF, Word, PPTX ...) extraction and parse API
Extract one time password (OTP) secrets from QR codes
Read and extract text and other content from PDFs in C#
A pure-python PDF library capable of splitting, merging, cropping
JavaScript OCR and text extraction for images and PDFs
PDFsam, a desktop application to split, merge, mix, rotate PDF files
Comprehensive Gradio WebUI for audio processing
OCR model for complex documents with layout-aware structured outputs
OCR software, free and offline
A cross-platform software for text translation and recognition
A Node.js scraper for humans
Ksoup is a lightweight Kotlin Multiplatform library for parsing HTML
A simple native web interface that uses ChatTTS to synthesize text
A fast, helpful, and open-source document parser
Handwritten Text Recognition (HTR) system implemented with TensorFlow
Qwen's most powerful open-source image generation model
Open source semantic search and text analytics for large document sets
YAML templating tool that works on YAML structure instead of text
The Refactoring library based off the Refactoring book
Python bindings for MuPDF's rendering library.
Public opinion analysis system
Multi-tool for semantic search
CLI tool to extract (meta)data from PDF and manipulate PDF files
Create prompt-friendly codebase digests from any Git repository URL
Generate blog articles from video or audio