PyTorch extensions for fast R&D prototyping and Kaggle farming
simplejson is a simple, fast, extensible JSON encoder/decoder
Bringing BERT into modernity via both architecture changes and scaling
Unified Multimodal Understanding and Generation Models
Industrial-level controllable zero-shot text-to-speech system
This repository contains the official implementation of FastVLM
Collection of Gemma 3 variants that are trained for performance
A high quality MP3 encoder
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Open-source industrial-grade ASR models
PyTorch code and models for V-JEPA self-supervised learning from video
Self-supervised visual learning using momentum contrast in PyTorch
End-to-end speech processing toolkit
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
PyTorch code and models for VJEPA2 self-supervised learning from video
AV1 Image File Format Specification - ISO-BMFF/HEIF derivative
Fast multimodal LLM for real-time voice interaction and AI apps
Retrieval and Retrieval-augmented LLMs
Visual Causal Flow
Open-source image generative foundation model
Official inference repo for FLUX.2 models
An open-source toolkit for BigMac-style pipeline-parallel training
Video-based AI memory library. Store millions of text chunks in MP4
Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Accurate × Fast × Comprehensive