PyTorch extensions for fast R&D prototyping and Kaggle farming
Bringing BERT into modernity via both architecture changes and scaling
Unified Multimodal Understanding and Generation Models
Industrial-level controllable zero-shot text-to-speech system
This repository contains the official implementation of FastVLM
Collection of Gemma 3 variants that are trained for performance
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Open-source industrial-grade ASR models
PyTorch code and models for V-JEPA self-supervised learning from video
Self-supervised visual learning using momentum contrast in PyTorch
End-to-end speech processing toolkit
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
PyTorch code and models for VJEPA2 self-supervised learning from video
Fast inference engine for Transformer models
Visual Causal Flow
Video-based AI memory library. Store millions of text chunks in MP4
Open-source image generative foundation model
Moonshot's most powerful AI model
Official inference repo for FLUX.2 models
Native MLX runtime for Laya typed decision models
A simple but complete full-attention transformer
An open-source toolkit for BigMac-style pipeline-parallel training
Accurate × Fast × Comprehensive
Qwen2.5-VL is the multimodal large language model series
C++ Implementation of PyTorch Tutorials for Everyone