Code for running inference with the SAM 3D Body Model 3DB
Personalize Any Characters with a Scalable Diffusion Transformer
Open-source image generative foundation model
Implementation of "MobileCLIP" CVPR 2024
From Images to High-Fidelity 3D Assets
Models for object and human mesh reconstruction
Qwen-Image is a powerful image generation foundation model
Generate Any 3D Scene in Seconds
Multimodal Diffusion with Representation Alignment
Unified Multimodal Understanding and Generation Models
Repo of Qwen2-Audio chat & pretrained large audio language model
Foundational Models for State-of-the-Art Speech and Text Translation
RGBD video generation model conditioned on camera input
code for Mesh R-CNN, ICCV 2019
Capable of understanding text, audio, vision, video
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
Powerful open source image generation model
ChatGLM-6B: An Open Bilingual Dialogue Language Model
A state-of-the-art open visual language model
Learning to Act by Watching Unlabeled Online Videos