ICLR2024 Spotlight: curation/training code, metadata, distribution
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
A Customizable Image-to-Video Model based on HunyuanVideo
An Open Real-time Video-Language Interaction System
Recovering the Visual Space from Any Views
Hunyuan Translation Model Version 1.5
VMZ: Model Zoo for Video Modeling
Tool for exploring and debugging transformer model behaviors
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
Miso TTS is an 8 billion, highly emotive text-to-speech model
Code for running inference with the SAM 3D Body Model 3DB
Unified Multimodal Understanding and Generation Models
Uncommon Objects in 3D dataset
Qwen3-omni is a natively end-to-end, omni-modal LLM
FAIR Sequence Modeling Toolkit 2
GPT4V-level open-source multi-modal model based on Llama3-8B
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
AI PPT Track Terminator, the strongest PPT Skill ever
Netease Youdao's open-source embedding and reranker models
An Efficient Agentic Model for Computer Use
Fast and Universal 3D reconstruction model for versatile tasks
PyTorch code and models for the DINOv2 self-supervised learning
Capable of understanding text, audio, vision, video
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions