CLIP, Predict the most relevant text snippet given an image
Reference PyTorch implementation and models for DINOv3
Implementation of "MobileCLIP" CVPR 2024
PyTorch code and models for the DINOv2 self-supervised learning
PyTorch implementation of MAE
CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B
CLIP model fine-tuned for zero-shot fashion product classification
Multimodal Transformer for document image understanding and layout
Small 3B-base multimodal model ideal for custom AI on edge hardware
Versatile 8B-base multimodal LLM, flexible foundation for custom AI