Motion-controllable Video Generation via Latent Trajectory Guidance
Large Multimodal Models for Video Understanding and Editing
RGBD video generation model conditioned on camera input
Implementation of Phenaki Video, which uses Mask GIT
Local-first AI image generation. No ComfyUI, no Python setup.
Open-source AI video pipeline, fully automated with MCP
A Customizable Image-to-Video Model based on HunyuanVideo
Overcoming Data Limitations for High-Quality Video Diffusion Models
Implementation of Video Diffusion Models
Implementation of Make-A-Video, new SOTA text to video generator
Implementation of Recurrent Interface Network (RIN)
CLIP + FFT/DWT/RGB = text to image/video
Multimodal AI Story Teller, built with Stable Diffusion, GPT, etc.
A walk along memory lane
Implementation of NÜWA, attention network for text to video synthesis
Implementation of NWT, audio-to-video generation, in Pytorch
Software tool that converts text to video for more engaging experience
The leading software for creating deepfakes
DCVGAN: Depth Conditional Video Generation, ICIP 2019.