Accurate × Fast × Comprehensive
Recovering the Visual Space from Any Views
Clean and efficient FP8 GEMM kernels with fine-grained scaling
Designed for text embedding and ranking tasks
Long-form streaming TTS system for multi-speaker dialogue generation
High-Fidelity and Controllable Generation of Textured 3D Assets
Large Multimodal Models for Video Understanding and Editing
Implementation of the Surya Foundation Model for Heliophysics
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Visual Causal Flow
Repo of Qwen2-Audio chat & pretrained large audio language model
Plugin and skin collection for DeepSeek Harness (DSH) Web UI
Block Diffusion for Ultra-Fast Speculative Decoding
MiniMax-M2, a model built for Max coding & agentic workflows
State of the art LLM and coding model
Proxy that exposes Antigravity provided claude / gemini models
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Claude Code image, a one-stop open source transit service
Repo for SeedVR2 & SeedVR
Collection of Gemma 3 variants that are trained for performance
The official PyTorch implementation of Google's Gemma models
A 0.1B Omni model trained from scratch
A Powerful Native Multimodal Model for Image Generation
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning