High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Industrial-level controllable zero-shot text-to-speech system
A Powerful Native Multimodal Model for Image Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
Controllable & emotion-expressive zero-shot TTS
Infinite Worlds with Versatile Interactions
HY-Motion model for 3D character animation generation
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Sharp Monocular Metric Depth in Less Than a Second
scikit-learn compatible tabular foundation model
MOSS‑TTS Family open‑source speech and sound generation model
Implementation of "MobileCLIP" CVPR 2024
CLIP, Predict the most relevant text snippet given an image
Language modeling in a sentence representation space
Pretrained time-series foundation model developed by Google Research
Production-tested AI infrastructure tools
ICLR2024 Spotlight: curation/training code, metadata, distribution
LLM-based Reinforcement Learning audio edit model
StudioOllamaUI is a local, portable interface for Ollama
Chat & pretrained large audio language model proposed by Alibaba Cloud
800,000 step-level correctness labels on LLM solutions to MATH problem
PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)
A library for Multilingual Unsupervised or Supervised word Embeddings
Open reasoning model for agentic coding and tool workflows
Trillion-parameter MoE model for coding and million-token reasoning