Make any agent harness multimodal-native
OpenRecall is a fully open-source, privacy-first alternative
A Repo For Document AI
OCR model for complex documents with layout-aware structured outputs
Document content and metadata extraction microservice
Structured data extraction and instruction calling with ML, LLM
A community-supported supercharged version of paperless
LLM inference server with continuous batching & SSD caching
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Let your agent control your phone
Pycorrector is a toolkit for text error correction
AI tool for automating desktop tasks via natural language input
An on-premises, OCR-free unstructured data extraction
Handwritten Text Recognition (HTR) system implemented with TensorFlow
In-depth tutorials on LLMs, RAGs and real-world AI agent applications
Qwen3-omni is a natively end-to-end, omni-modal LLM
Visual Automation IDE — automate anything you see on screen
Visual desktop automation, OCR, JSON & web automation toolkit
A Python application to add watermarks (text or image) to PDF files
Desktop research workspace for PDFs, notes, citations, bibliographies.
Ferramenta de Tarjamento de Dados Pessoais e Sigilosos
Vision utilities for web interaction agents
FaceOnLive Open KYC: Streamlining Identity Verification with AI
Implementation of Nougat Neural Optical Understanding