Stable Virtual Camera: Generative View Synthesis with Diffusion Models
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Large Multimodal Models for Video Understanding and Editing
RGBD video generation model conditioned on camera input
Netease Youdao's open-source embedding and reranker models
An Efficient Agentic Model for Computer Use
Revolutionizing Database Interactions with Private LLM Technology
Pokee Deep Research Model Open Source Repo
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Renderer for the harmony response format to be used with gpt-oss
Generating Immersive, Explorable, and Interactive 3D Worlds
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Miso TTS is an 8 billion, highly emotive text-to-speech model
The official PyTorch implementation of Google's Gemma models
Qwen3-omni is a natively end-to-end, omni-modal LLM
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Open Source Speech Language Model
Foundation model for image generation
Hunyuan Translation Model Version 1.5