Expressive Portrait Image Animation for Live Streaming
Diffusion Transformer with Fine-Grained Chinese Understanding
UniMate: One Unified Model to Animate Diverse Skeletons
Repo for SeedVR2 & SeedVR
Tokenizer-Free TTS for Multilingual Speech Generation
Multimodal Diffusion with Representation Alignment
Open-source multi-speaker long-form text-to-speech model
HY-Motion model for 3D character animation generation
RGBD video generation model conditioned on camera input
Cosmos-RL is a flexible and scalable Reinforcement Learning framework
A SOTA open-source image editing model
ComfyUI wrapper nodes for WanVideo and related models
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Official PyTorch Implementation
Personalize Any Characters with a Scalable Diffusion Transformer
Official inference repo for FLUX.1 models
Project Lyra: Open Generative 3D World Models
Code and models for ICML 2024 paper, NExT-GPT
Stable Diffusion web UI
Inference script for Oasis 500M
Deep learning framework
State-of-the-art (SoTA) text-to-video pre-trained model
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Pytorch Distributed native training library for LLMs/VLMs
InvokeAI is a leading creative engine for Stable Diffusion models