CLIP, Predict the most relevant text snippet given an image
Easily compute clip embeddings and build a clip retrieval system
An open source implementation of CLIP
ICLR2024 Spotlight: curation/training code, metadata, distribution
Automatically translates the text of a video based on a subtitle file
Suite of reference architectures for building GPU-accelerated vision
Instant voice cloning by MIT and MyShell. Audio foundation model
The most powerful and modular diffusion model GUI, api and backend
TorchMultimodal is a PyTorch library
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences
Cross platform GUI tool for downloading videos from Bilibili sites
Stable Diffusion web UI
Ableton Live Model Context Protocol Integration
High-Quality Voice Cloning TTS for 600+ Languages
LTX-Video Support for ComfyUI
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Interface for OuteTTS models
Gracefully face hCaptcha challenge with multimodal llms
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Generating Immersive, Explorable, and Interactive 3D Worlds
A Python library for audio data augmentation
Multi-Modal Neural Networks for Semantic Search, based on Mid-Fusion
Tensor search for humans
The data structure for multimodal data
Implementation of Imagen, Google's Text-to-Image Neural Network