A collection of tools, libraries, and tests for Vulkan shader
HunyuanVideo: A Systematic Framework For Large Video Generation Model
A 0.1B Omni model trained from scratch
LLM plugin providing access to models running on an Ollama server
All-in-one native macOS AI chat application
"Big Model" trains a visual multimodal VLM with 26M parameters
Simplifies the local serving of AI models from any source
Implementation of "MobileCLIP" CVPR 2024
NLP Cloud serves high performance pre-trained or custom models for NER
The data structure for multimodal data
The repository provides code for running inference with SAM 2
Accurate × Fast × Comprehensive
Convert (animated) stickers to/from WhatsApp, Telegram, Signal
Advanced AI Explainability for computer vision
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Sharp Monocular View Synthesis in Less Than a Second
Official implementation of DreamCraft3D
Build cross-modal and multimodal applications on the cloud
Datasets, transforms and models specific to Computer Vision
Native InstantID support for ComfyUI
An extensive node suite that enables ComfyUI to process 3D inputs
Make any agent harness multimodal-native
Extract one time password (OTP) secrets from QR codes
Qwen3-omni is a natively end-to-end, omni-modal LLM