High-Fidelity and Controllable Generation of Textured 3D Assets
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A series of math-specific large language models of our Qwen2 series
Generating Immersive, Explorable, and Interactive 3D Worlds
Qwen2.5-VL is the multimodal large language model series
MOSS‑TTS Family open‑source speech and sound generation model
Bidirectional token-classification model for identifiable info
Genome modeling and design across all domains of life
Project Lyra: Open Generative 3D World Models
Achieving 3+ generation speedup on reasoning tasks
Ultra-Efficient LLMs on End Device
Pretrained time-series foundation model developed by Google Research
Fast and Universal 3D reconstruction model for versatile tasks
4M: Massively Multimodal Masked Modeling
ICLR2024 Spotlight: curation/training code, metadata, distribution
A PyTorch library for implementing flow matching algorithms
Official implementation of DreamCraft3D
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Repo of Qwen2-Audio chat & pretrained large audio language model
New family of code large language models (LLMs)