Parse files for optimal RAG
Multilingual sentence & image embeddings with BERT
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
ComfyUI wrapper nodes for WanVideo and related models
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Unified Multimodal Understanding and Generation Models
Fast-stable-diffusion + DreamBooth
"Big Model" trains a visual multimodal VLM with 26M parameters
Implementation of "MobileCLIP" CVPR 2024
Open source personal AI Assistant for Linux, Windows and Mac
An AI-agent skill that turns Markdown into paste-ready WeChat article
Multimodal embedding and reranking models built on Qwen3-VL
CLI tool to extract (meta)data from PDF and manipulate PDF files
Foundational video generation model with 13.6B parameters
21 Lessons, Get Started Building with Generative AI
Stable Diffusion built-in to Blender
Generate Any 3D Scene in Seconds
Official Python inference and LoRA trainer package
Pretrained model hub for Keras 3
Sample code and notebooks for Generative AI on Google Cloud
Open-Sora: Democratizing Efficient Video Production for All
Phi-3.5 for Mac: Locally-run Vision and Language Models
Extract one time password (OTP) secrets from QR codes
The data structure for multimodal data
Windrecorder is a memory search app by records everything