A Unified Framework for Text-to-3D and Image-to-3D Generation
Unifying 3D Mesh Generation with Language Models
Generating Immersive, Explorable, and Interactive 3D Worlds
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
A Multi-Modal World Model for Reconstructing, Generating, Simulation
HY-Motion model for 3D character animation generation
A python parametric CAD scripting framework based on OCCT
A text-to-speech, speech-to-text and speech-to-speech library
Generate Any 3D Scene in Seconds
Official implementation of DreamCraft3D
State-of-the-art (SoTA) text-to-video pre-trained model
[CVPR 2026 Oral] VGGT Omega
Open-Source Dual-Arm Mobile Robot with Motorized Lift
Make any agent harness multimodal-native
Framework for building AI-powered interactive digital humans and agent
A Python toolbox for gaining geometric insights
The data structure for multimodal data
A Systematic Framework for Interactive World Modeling
Framework for building neural networks
State-of-the-art diffusion models for image and audio generation
Towards Studio-Grade Character Animation via In-Context Learning of 3D
Circuit diagrams and firmware source code for Gboard DIY keyboards
Build cross-modal and multimodal applications on the cloud
2D & 3D TeX-Aware Vector Graphics Language
Implementation of Make-A-Video, new SOTA text to video generator