ExtractThinker is a Document Intelligence library for LLMs
Structured data extraction and instruction calling with ML, LLM
No-code LLM Platform to launch APIs and ETL Pipelines
ContextGem: Effortless LLM extraction from documents
A high-quality tool for convert PDF to Markdown and JSON
Document content and metadata extraction microservice
Document (PDF, Word, PPTX ...) extraction and parse API
RAGFlow is an open-source RAG (Retrieval-Augmented Generation) engine
Synthetic data curation for post-training and data extraction
Claude Code skill for generating production-quality SVG+PNG technical
End-to-end pipeline converting generative videos
Superlinked is a Python framework for AI Engineers
An on-premises, OCR-free unstructured data extraction
kaldi-asr/kaldi is the official location of the Kaldi project
Online machine learning in Python
AI video generator optimized for low VRAM and older GPUs use
PyTorch code and models for the DINOv2 self-supervised learning
A Simple and Universal Swarm Intelligence Engine
From Paper to Presentation in One Click
OCR expert VLM powered by Hunyuan's native multimodal architecture
Open-source evaluation toolkit of large multi-modality models (LMMs)
End-to-end speech processing toolkit
Tools to build web AI agents that can authenticate
File Parser optimised for LLM Ingestion with no loss
Get your documents ready for gen AI