Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Foundation Models for Time Series
Hackable and optimized Transformers building blocks
RGBD video generation model conditioned on camera input
Recovering the Visual Space from Any Views
Miso TTS is an 8 billion, highly emotive text-to-speech model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Unified Multimodal Understanding and Generation Models
Open-source large language model family from Tencent Hunyuan
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions
Tiny vision language model
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Implementation of the Surya Foundation Model for Heliophysics
scikit-learn compatible tabular foundation model
Audio foundation model excelling in audio understanding
Code for running inference with the SAM 3D Body Model 3DB
An Open Real-time Video-Language Interaction System
Video Object and Interaction Deletion
Qwen3-ASR is an open-source series of ASR models
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Fast-stable-diffusion + DreamBooth
Accurate × Fast × Comprehensive
An experimental version of DeepSeek model
1B text generation model based on the HRM architecture