A Multi-Modal World Model for Reconstructing, Generating, Simulation
A SOTA open-source image editing model
Open-source framework for intelligent speech interaction
Unified Multimodal Understanding and Generation Models
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
PyTorch code and models for the DINOv2 self-supervised learning
Official implementation of DreamCraft3D
Open-source industrial-grade ASR models
Official implementation of Watermark Anything with Localized Messages
Multimodal-Driven Architecture for Customized Video Generation
Personalize Any Characters with a Scalable Diffusion Transformer
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
4M: Massively Multimodal Masked Modeling
ICLR2024 Spotlight: curation/training code, metadata, distribution
CogView4, CogView3-Plus and CogView3(ECCV 2024)
A Customizable Image-to-Video Model based on HunyuanVideo
High-Fidelity and Controllable Generation of Textured 3D Assets
RGBD video generation model conditioned on camera input
Controllable & emotion-expressive zero-shot TTS
Large-language-model & vision-language-model based on Linear Attention
ChatGPT interface with better UI
Chat & pretrained large audio language model proposed by Alibaba Cloud
Real-time behaviour synthesis with MuJoCo, using Predictive Control
Example Discord bot written in Python that uses the completions API
Code for the paper Hybrid Spectrogram and Waveform Source Separation