pdf-inspector
Fast Rust library for PDF inspection, classification
...It distinguishes text-based, scanned, image-based, and mixed documents while returning confidence scores and pages that may need OCR. Its position-aware parser preserves font data, coordinates, reading order, and multi-column layouts. The converter produces clean Markdown with headings, lists, code blocks, links, page breaks, and formatted tables. It supports CID fonts, several encodings, right-to-left text, and automatic detection of broken font mappings. The same core is available through Rust, Python, Node.js, browser WebAssembly, and command-line interfaces. It is designed for local document pipelines that need fast routing before using slower OCR services.