A Multi-Modal World Model for Reconstructing, Generating, Simulation
A SOTA open-source image editing model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open-source industrial-grade ASR models
Official implementation of Watermark Anything with Localized Messages
Video understanding codebase from FAIR for reproducing video models
Personalize Any Characters with a Scalable Diffusion Transformer
Controllable & emotion-expressive zero-shot TTS
Unified Multimodal Understanding and Generation Models
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Generating Immersive, Explorable, and Interactive 3D Worlds
Achieving 3+ generation speedup on reasoning tasks
4M: Massively Multimodal Masked Modeling
ICLR2024 Spotlight: curation/training code, metadata, distribution
Official implementation of DreamCraft3D
A Customizable Image-to-Video Model based on HunyuanVideo
High-Fidelity and Controllable Generation of Textured 3D Assets
RGBD video generation model conditioned on camera input
Large-language-model & vision-language-model based on Linear Attention
ChatGPT interface with better UI
Chat & pretrained large audio language model proposed by Alibaba Cloud
Real-time behaviour synthesis with MuJoCo, using Predictive Control
Example Discord bot written in Python that uses the completions API
Code for the paper Hybrid Spectrogram and Waveform Source Separation