Multimodal-Driven Architecture for Customized Video Generation
Official Python inference and LoRA trainer package
Pushing the Frontier of Long Audio-Visual Generation
Foundational video generation model with 13.6B parameters
A python tool that uses GPT-4, FFmpeg, and OpenCV
HunyuanVideo: A Systematic Framework For Large Video Generation Model
Large Multimodal Models for Video Understanding and Editing
Multimodal AI Story Teller, built with Stable Diffusion, GPT, etc.
A walk along memory lane
Implementation of NÜWA, attention network for text to video synthesis