Document (PDF, Word, PPTX ...) extraction and parse API
A Java library used to read and extract data from NFC EMV credit cards
Extrect selected entries from LDIF files like grep
Extract public Instagram account information from usernames
A pure-python PDF library capable of splitting, merging, cropping
Extract structured data from webpages using LLM-powered scraping
Crawl a website starting from a URL, find relevant pages
High-performance Rust web crawler and scraper for large-scale data
PDFCraft is a free, privacy-focused PDF toolkit
AI-ready web crawler that extracts and structures website content
Saves Discord chat logs to a file
Python crawler for collecting and downloading Sina Weibo user data
An on-premises, OCR-free unstructured data extraction
OCR model for complex documents with layout-aware structured outputs
Fast and efficient unstructured data extraction
Asyncio-based Python framework for building fast web crawling spiders
To extract main article from given URL with Node.js
Library for reducing tail latency in RAM reads
Open source semantic search and text analytics for large document sets
Your clothes, extracted and organized with gpt-image
Progressive PHP web crawler framework with jQuery-like DOM parsing
A system for agentic LLM-powered data processing and ETL
The undetected self-hosted browser automation platform
Component of phpDocumentor provides a DocBlock parser
Contexts Optical Compression