A Unified Framework for Text-to-3D and Image-to-3D Generation
Unifying 3D Mesh Generation with Language Models
Generating Immersive, Explorable, and Interactive 3D Worlds
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
A Multi-Modal World Model for Reconstructing, Generating, Simulation
HY-Motion model for 3D character animation generation
Generate Any 3D Scene in Seconds
A text-to-speech, speech-to-text and speech-to-speech library
Official implementation of DreamCraft3D
State-of-the-art (SoTA) text-to-video pre-trained model
Make any agent harness multimodal-native
Framework for building AI-powered interactive digital humans and agent
The data structure for multimodal data
A Systematic Framework for Interactive World Modeling
Framework for building neural networks
State-of-the-art diffusion models for image and audio generation
Build cross-modal and multimodal applications on the cloud
Implementation of Make-A-Video, new SOTA text to video generator
Implementation of Video Diffusion Models
Generate 3D objects conditioned on text or images
Framework that is dedicated to making neural data processing
CLIP + FFT/DWT/RGB = text to image/video
Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion
A walk along memory lane
Point cloud diffusion for 3D model synthesis