Native and Compact Structured Latents for 3D Generation
Reference PyTorch implementation and models for DINOv3
Multimodal-Driven Architecture for Customized Video Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
AI PPT Track Terminator, the strongest PPT Skill ever
Lets make video diffusion practical
Personalize Any Characters with a Scalable Diffusion Transformer
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Fast-stable-diffusion + DreamBooth
Official implementation of Watermark Anything with Localized Messages
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Diffusion Transformer with Fine-Grained Chinese Understanding
GPT4V-level open-source multi-modal model based on Llama3-8B
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Capable of understanding text, audio, vision, video
Sharp Monocular Metric Depth in Less Than a Second
High-Resolution Image Synthesis with Latent Diffusion Models
Codex plugin that turns attached object images into code-only
Recovering the Visual Space from Any Views
Contexts Optical Compression
Official Python inference and LoRA trainer package
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Unified Multimodal Understanding and Generation Models
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Phi-3.5 for Mac: Locally-run Vision and Language Models