Pretrained time-series foundation model developed by Google Research
General-purpose image editing model that delivers high-fidelity
Fast and Universal 3D reconstruction model for versatile tasks
Foundation Models for Time Series
ICLR2024 Spotlight: curation/training code, metadata, distribution
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
New family of code large language models (LLMs)
FAIR Sequence Modeling Toolkit 2
Language modeling in a sentence representation space
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
RGBD video generation model conditioned on camera input
Chinese and English multimodal conversational language model
Tooling for the Common Objects In 3D dataset
High-Resolution Image Synthesis with Latent Diffusion Models
Qwen2.5-Coder is the code version of Qwen2.5, the large language model
A state-of-the-art open visual language model
Chat & pretrained large audio language model proposed by Alibaba Cloud
Chat & pretrained large vision language model
Towards Real-World Vision-Language Understanding
Real-time behaviour synthesis with MuJoCo, using Predictive Control
Official DeiT repository
Dataset of GPT-2 outputs for research in detection, biases, and more
Code for the paper Hybrid Spectrogram and Waveform Source Separation
This repository contains the official implementation of research