A collection of high-quality models for the MuJoCo physics engine
GLM-4-Voice | End-to-End Chinese-English Conversational Model
An Efficient Agentic Model for Computer Use
A 0.1B Omni model trained from scratch
26m function call model that runs on incredibly small devices
High-resolution models for human tasks
Long-form streaming TTS system for multi-speaker dialogue generation
Open-source deep-learning framework
Inference script for Oasis 500M
Advancing Open-source World Models
A Systematic Framework for Interactive World Modeling
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
OCR expert VLM powered by Hunyuan's native multimodal architecture
The ChatGPT Retrieval Plugin lets you easily find personal documents
Real-time behaviour synthesis with MuJoCo, using Predictive Control
Official repo for consistency models
Repo for external large-scale work
A GUI tool for generating subtitle from videos, generating srt files
PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)
Reference implementation of the Transformer architecture optimized
GLIDE: a diffusion-based text-conditional image synthesis model
Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)
Dual LSTM Encoder for Dialog Response Generation
Open language model developed by NVIDIA as part of Nemotron-3 family
Open-source code agent designed for Lean 4