JoyAI-Video-Edit
Real-Time Open-Ended Video Editing with Autoregressive Diffusion
...Natural-language instructions can drive subject edits, local changes, background replacement, style transfer, motion changes, and reference-guided transformations. Its architecture combines a multimodal condition encoder, causal video VAE, and a 16B-parameter multimodal diffusion transformer. Autoregressive diffusion and bounded KV-state inference are used to keep computation stable across long streams. The released deployment reaches high-throughput 720p editing and also supports real-time operation on selected consumer GPUs. Updated checkpoints improve identity preservation, reference conditioning, and temporal consistency for reference-image-guided editing.