Text and image to video generation: CogVideoX and CogVideo
One-click generation of various gameplay, no prompt words required
Collection of Gemma 3 variants that are trained for performance
Code for running inference with the SAM 3D Body Model 3DB
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Native and Compact Structured Latents for 3D Generation
PyTorch implementation of JiT
Reference PyTorch implementation and models for DINOv3
Multimodal-Driven Architecture for Customized Video Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
AI PPT Track Terminator, the strongest PPT Skill ever
Lets make video diffusion practical
Personalize Any Characters with a Scalable Diffusion Transformer
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
DeepSeek Harness (DSH) Web
Fast-stable-diffusion + DreamBooth
Official implementation of Watermark Anything with Localized Messages
Diffusion Transformer with Fine-Grained Chinese Understanding
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Plugin and skin collection for DeepSeek Harness (DSH) Web UI
GPT4V-level open-source multi-modal model based on Llama3-8B
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Capable of understanding text, audio, vision, video
High-Resolution Image Synthesis with Latent Diffusion Models
Sharp Monocular Metric Depth in Less Than a Second