Document (PDF, Word, PPTX ...) extraction and parse API
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
Zero-copy PDF text extraction library written in Zig
Python & command-line tool to gather text on the Web
CLI tool to extract (meta)data from PDF and manipulate PDF files
Structured data extraction and instruction calling with ML, LLM
NLP Cloud serves high performance pre-trained or custom models for NER
Document content and metadata extraction microservice
Open source NLP guide with models, methods, and real use cases
OCR software, free and offline
Python module for parsing semi-structured text into python tables
Knowledge Graph Generation from Any Text
A Family of Open Sourced Music Foundation Models
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Edit PDF files with Nano Banana
A simple tool for reading in poorly redacted documents
RAG-Anything: All-in-One RAG Framework
Python binding to the Apache Tika™ REST services
Accurate × Fast × Comprehensive
Open source healthcare AI
Python library for scraping and analyzing online news articles easily
A high-quality PDF to Markdown tool based on large language model
Reading book source
OCR expert VLM powered by Hunyuan's native multimodal architecture
Open source OSINT tool for gathering data on emails, phones, and IPs