pdf-inspector is a fast Rust library for classifying PDFs and extracting structured text without OCR. It distinguishes text-based, scanned, image-based, and mixed documents while returning confidence scores and pages that may need OCR. Its position-aware parser preserves font data, coordinates, reading order, and multi-column layouts. The converter produces clean Markdown with headings, lists, code blocks, links, page breaks, and formatted tables. It supports CID fonts, several encodings, right-to-left text, and automatic detection of broken font mappings. The same core is available through Rust, Python, Node.js, browser WebAssembly, and command-line interfaces. It is designed for local document pipelines that need fast routing before using slower OCR services.

Features

  • Smart PDF classification with confidence scores
  • Per-page OCR routing recommendations
  • Position-aware text and layout extraction
  • Structured Markdown and table conversion
  • CID font and multiple encoding support
  • Rust, Python, Node.js, WebAssembly, and CLI access

Project Samples

Project Activity

See All Activity >

Categories

Libraries

License

MIT License

Follow pdf-inspector

pdf-inspector Web Site

Other Useful Business Software
Demo Series - Small Business Backup By Veeam Icon
Demo Series - Small Business Backup By Veeam

Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
Watch Demo Series
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of pdf-inspector!

Additional Project Details

Programming Language

Rust

Related Categories

Rust Libraries

Registered

2026-08-04