pdf-inspector is a fast Rust library for classifying PDFs and extracting structured text without OCR. It distinguishes text-based, scanned, image-based, and mixed documents while returning confidence scores and pages that may need OCR. Its position-aware parser preserves font data, coordinates, reading order, and multi-column layouts. The converter produces clean Markdown with headings, lists, code blocks, links, page breaks, and formatted tables. It supports CID fonts, several encodings, right-to-left text, and automatic detection of broken font mappings. The same core is available through Rust, Python, Node.js, browser WebAssembly, and command-line interfaces. It is designed for local document pipelines that need fast routing before using slower OCR services.

Features

  • Smart PDF classification with confidence scores
  • Per-page OCR routing recommendations
  • Position-aware text and layout extraction
  • Structured Markdown and table conversion
  • CID font and multiple encoding support
  • Rust, Python, Node.js, WebAssembly, and CLI access

Project Samples

Project Activity

See All Activity >

Categories

Libraries

License

MIT License

Follow pdf-inspector

pdf-inspector Web Site

Other Useful Business Software
Ship Agents Faster Icon
Ship Agents Faster

Transform your applications and workflows into powerful agentic systems at global scale.

Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of pdf-inspector!

Additional Project Details

Programming Language

Rust

Related Categories

Rust Libraries

Registered

2026-08-04