Contexts Optical Compression
Accurate × Fast × Comprehensive
PDF to Markdown with vision models
Visual Causal Flow
Awesome multilingual OCR toolkits based on PaddlePaddle
Convert AI papers to GUI
Use LLMs and LLM Vision (OCR) to handle paperless-ngx
A framework to enable multimodal models to operate a computer
In-depth tutorials on LLMs, RAGs and real-world AI agent applications
Declarative way to run AI models in React Native on device
PDF scientific paper translation with preserved formats
Enhances Tesseract OCR output using LLMs (local or API)
Screenshots, word marking, OCR, AI, translation software
OCR expert VLM powered by Hunyuan's native multimodal architecture
Make any agent harness multimodal-native
Get your documents ready for gen AI
Readest is a modern, feature-rich ebook reader
PDF Parser for AI-ready data. Automate PDF accessibility
A mouse-themed, offline Windows file converter
A Repo For Document AI
Qwen3-VL, the multimodal large language model series by Alibaba Cloud
Streamline your life using PromptingTools.jl
OpenRecall is a fully open-source, privacy-first alternative
Doctor Dok is an AI based medical data framework
Let your agent control your phone