JoyAI-Echo
Pushing the Frontier of Long Audio-Visual Generation
...It is designed to create minute-level, multi-shot video stories from structured prompts while preserving continuity across scenes. The system uses a paired cross-modal memory bank to maintain visual identity and voice consistency over longer sequences. It also uses a distilled DMD generator to reduce inference cost and improve generation speed compared with heavier multi-step pipelines. JoyAI-Echo focuses on text-to-video and multi-shot long-video generation, while image-to-video support is not part of the current release scope. It is most useful for research and experimental video workflows that need synchronized audio, coherent characters, and editable story-level generation.