CLIP, Predict the most relevant text snippet given an image
New family of code large language models (LLMs)
4M: Massively Multimodal Masked Modeling
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Python inference and LoRA trainer package for the LTX-2 audio–video
PyTorch code and models for the DINOv2 self-supervised learning
A Powerful Native Multimodal Model for Image Generation
An experimental version of DeepSeek model
Block Diffusion for Ultra-Fast Speculative Decoding
Pretrained time-series foundation model developed by Google Research
ICLR2024 Spotlight: curation/training code, metadata, distribution
LLM-based Reinforcement Learning audio edit model
ChatGPT interface with better UI
Open-source, high-performance Mixture-of-Experts large language model
The ChatGPT Retrieval Plugin lets you easily find personal documents
Official code for Style Aligned Image Generation via Shared Attention
Towards Robust Blind Face Restoration with Codebook Lookup Transformer
Fine-tuning ChatGLM-6B with PEFT
A minimal PyTorch re-implementation of the OpenAI GPT
Reference implementation of the Transformer architecture optimized
Large-scale autoregressive pixel model for image generation by OpenAI
A library for Multilingual Unsupervised or Supervised word Embeddings