CLI tool to extract (meta)data from PDF and manipulate PDF files
ExtractThinker is a Document Intelligence library for LLMs
Structured data extraction and instruction calling with ML, LLM
No-code LLM Platform to launch APIs and ETL Pipelines
ContextGem: Effortless LLM extraction from documents
Turn any technical book PDF into a Claude Code skill
AI-ready web crawler that extracts and structures website content
A high-quality tool for convert PDF to Markdown and JSON
Web Robotics Process Automation Tool
Document content and metadata extraction microservice
Zero-copy PDF text extraction library written in Zig
Document (PDF, Word, PPTX ...) extraction and parse API
RAGFlow is an open-source RAG (Retrieval-Augmented Generation) engine
Synthetic data curation for post-training and data extraction
Asyncio-based Python framework for building fast web crawling spiders
Claude Code skill for generating production-quality SVG+PNG technical
Python & command-line tool to gather text on the Web
Burp Suite extension for JavaScript static analysis
Python module for parsing semi-structured text into python tables
Python tool for crawling and extracting structured data from news site
Python3 web crawler practice
End-to-end pipeline converting generative videos
TikTok releases/likes/compilations/live streams/videos/atlases/music
Superlinked is a Python framework for AI Engineers
Edit PDF files with Nano Banana