CLI tool to extract (meta)data from PDF and manipulate PDF files
ExtractThinker is a Document Intelligence library for LLMs
Fast, local-first web content extraction for LLMs
Structured data extraction and instruction calling with ML, LLM
Clean network diagrams, One-time setup, zero upkeep
Unreal Engine Archives Explorer
MD/.JSON Document OCR and structured data extraction API
Turn entire websites into LLM-ready markdown or structured data
No-code LLM Platform to launch APIs and ETL Pipelines
Crawl a website starting from a URL, find relevant pages
AI-first Ruby framework for building fast, flexible web scraping spide
PDF Parser for AI-ready data. Automate PDF accessibility
Model Context Protocol server that integrates AgentQL's data
Fast and efficient unstructured data extraction
Flexible Node.js AI-assisted crawler library
Turn any technical book PDF into a Claude Code skill
Automatic extraction of relevant features from time series
ContextGem: Effortless LLM extraction from documents
AI-ready web crawler that extracts and structures website content
BlockArrays for Julia
Declarative web scraping
A high-quality tool for convert PDF to Markdown and JSON
Extract and convert data from any document, images, pdfs, word doc
Web Robotics Process Automation Tool
Library for extracting streaming site data without official APIs