LingBot-Video is a large-scale mixture-of-experts video generation model focused on embodied intelligence. It is designed to connect video synthesis with physical-world understanding instead of generating only visually appealing clips. The project includes dense and MoE model variants for text-to-image, text-to-video, and text-image-to-video workflows. Its training combines large-scale web video data with more than 70,000 hours of embodied data. A multi-reward system emphasizes aesthetics, physical rationality, and task completion. The repository includes models, inference code, prompt rewriting tools, refiner workflows, and single-GPU or multi-GPU scripts.
Features
- MoE video generation model
- Dense 1.3B and MoE 30B-A3B variants
- Text-to-image, text-to-video, and TI2V support
- Prompt rewriter and LoRA adapter
- Base generation and refiner workflows
- Single-GPU and multi-GPU inference scripts
Categories
AI ModelsLicense
Apache License V2.0Follow LingBot-Video
Other Useful Business Software
Ship Agents Faster
Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of LingBot-Video!