A high-quality tool for convert PDF to Markdown and JSON
Get your documents ready for gen AI
An Open-Source Toolkit for General-OCR Research and Applications
Multilingual Document Layout Parsing in a Single Vision-Language Model
Contexts Optical Compression
An on-premises, OCR-free unstructured data extraction
Open source semantic search and text analytics for large document sets
OCR software, free and offline
A Repo For Document AI
Library for OCR-related tasks powered by Deep Learning
Canvas-based WYSIWYG rich text editor with advanced layout tools
Welcome the Era of One-shot Long-horizon Parsing
Enhances Tesseract OCR output using LLMs (local or API)
OCR model for complex documents with layout-aware structured outputs
Map location picker component for Android
The SILE Typesetter — Simon’s Improved Layout Engine
LaTeX template for undergraduate/graduate theses/dissertations
Open-Source Python3 tool for recognizing layouts, tables, and math
Accurate × Fast × Comprehensive
Assist in organizing your piles of documents
Extract and convert data from any document, images, pdfs, word doc
OCR expert VLM powered by Hunyuan's native multimodal architecture
Collabora Online is a collaborative online office suite
CLI tool to extract (meta)data from PDF and manipulate PDF files
Video translation and dubbing tool powered by LLMs