Command-line toolset for extracting text from files (documents, images, archives) into SQLite with OCR support.
Simple, expandable, one shell script only.

Features

  • Multi-format text extraction from 30+ file types including documents, spreadsheets, presentations, and archives
  • OCR (Optical Character Recognition) support for extracting text from images and scanned documents
  • Recursive archive processing - automatically extracts and processes files nested within ZIP, TAR, GZIP, and other archive formats
  • SQLite database integration - stores all extracted text in a searchable SQLite database for fast queries
  • Command-line interface - easy integration into scripts and automated workflows
  • Batch processing - process entire directories with a single command
  • Line-level granularity - extracts text with line numbers for precise referencing
  • Configurable OCR - supports multiple languages and quality settings
  • LibreOffice integration - uses headless LibreOffice for reliable document conversion
  • Lightweight - shell script implementation with minimal dependencies
  • Cross-platform - runs on Linux, macOS, and Windows (via WSL/Cygwin)
  • Transaction-safe database updates - uses SQL transactions for data integrity
  • Progress tracking - detailed output for monitoring extraction progress
  • Error handling - continues processing even if individual files fail

Project Activity

See All Activity >

Follow UniversalTextExtractor

UniversalTextExtractor Web Site

Other Useful Business Software
AI-powered service management for IT and enterprise teams Icon
AI-powered service management for IT and enterprise teams

Enterprise-grade ITSM, for every business

Give your IT, operations, and business teams the ability to deliver exceptional services—without the complexity. Maximize operational efficiency with refreshingly simple, AI-powered Freshservice.
Try it Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of UniversalTextExtractor!

Additional Project Details

Operating Systems

Linux, Windows

Intended Audience

Advanced End Users, System Administrators

User Interface

Command-line

Programming Language

C++, Unix Shell

Database Environment

SQLite

Related Categories

Unix Shell Search Software, Unix Shell Document Management System, Unix Shell File Tagging Software, C++ Search Software, C++ Document Management System, C++ File Tagging Software

Registered

2026-01-16