[CVPR 2026 Oral] VGGT Omega
Dealing with all unstructured data, such as reverse image search
Parse files for optimal RAG
Multilingual sentence & image embeddings with BERT
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Unified Multimodal Understanding and Generation Models
Fast-stable-diffusion + DreamBooth
PDFCraft is a free, privacy-focused PDF toolkit
Bridging framework for connecting arbitrary AI agents
"Big Model" trains a visual multimodal VLM with 26M parameters
Implementation of "MobileCLIP" CVPR 2024
Powerful yet simple to use screenshot software 🖥️ 📸
Lightweight Markdown app to help you write great sentences
An AI-agent skill that turns Markdown into paste-ready WeChat article
Multimodal embedding and reranking models built on Qwen3-VL
Open source personal AI Assistant for Linux, Windows and Mac
MOROS: Obscure Rust Operating System
Open-source framework for conversational voice AI agents
Ollama client that simplifies experimenting with LLMs
A minimal, accessible and SEO-friendly Astro blog theme
CLI tool to extract (meta)data from PDF and manipulate PDF files
Minimal PDF creation library
Foundational video generation model with 13.6B parameters
AI tool that generates custom presentations with real-time editing