pdf-inspector is a fast Rust library for classifying PDFs and extracting structured text without OCR. It distinguishes text-based, scanned, image-based, and mixed documents while returning confidence scores and pages that may need OCR. Its position-aware parser preserves font data, coordinates, reading order, and multi-column layouts. The converter produces clean Markdown with headings, lists, code blocks, links, page breaks, and formatted tables. It supports CID fonts, several encodings, right-to-left text, and automatic detection of broken font mappings. The same core is available through Rust, Python, Node.js, browser WebAssembly, and command-line interfaces. It is designed for local document pipelines that need fast routing before using slower OCR services.

Features

  • Smart PDF classification with confidence scores
  • Per-page OCR routing recommendations
  • Position-aware text and layout extraction
  • Structured Markdown and table conversion
  • CID font and multiple encoding support
  • Rust, Python, Node.js, WebAssembly, and CLI access

Project Samples

Project Activity

See All Activity >

Categories

Libraries

License

MIT License

Follow pdf-inspector

pdf-inspector Web Site

Other Useful Business Software
$300 Free Credits to Build on Google Cloud Icon
$300 Free Credits to Build on Google Cloud

New customers can spin up VMs, build with AI, and query data at no cost.

Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of pdf-inspector!

Additional Project Details

Programming Language

Rust

Related Categories

Rust Libraries

Registered

2026-08-04