Document (PDF, Word, PPTX ...) extraction and parse API
Extract one time password (OTP) secrets from QR codes
Read and extract text and other content from PDFs in C#
A pure-python PDF library capable of splitting, merging, cropping
JavaScript OCR and text extraction for images and PDFs
PDFsam, a desktop application to split, merge, mix, rotate PDF files
Comprehensive Gradio WebUI for audio processing
A cross-platform software for text translation and recognition
OCR software, free and offline
OCR model for complex documents with layout-aware structured outputs
Contexts Optical Compression
A simple native web interface that uses ChatTTS to synthesize text
Ksoup is a lightweight Kotlin Multiplatform library for parsing HTML
A fast, helpful, and open-source document parser
Handwritten Text Recognition (HTR) system implemented with TensorFlow
The Refactoring library based off the Refactoring book
Python bindings for MuPDF's rendering library.
Extract structured data from webpages using LLM-powered scraping
Open source semantic search and text analytics for large document sets
Generate blog articles from video or audio
A modular graph-based Retrieval-Augmented Generation (RAG) system
Create prompt-friendly codebase digests from any Git repository URL
CLI tool to extract (meta)data from PDF and manipulate PDF files
Automated translation solution for visual novels
PDFCraft is a free, privacy-focused PDF toolkit