A Pragmatic VLA Foundation Model
CLIP, Predict the most relevant text snippet given an image
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
Open-source framework for intelligent speech interaction
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
DeepMind model for tracking arbitrary points across videos & robotics
An Efficient Agentic Model for Computer Use
Revolutionizing Database Interactions with Private LLM Technology
scikit-learn compatible tabular foundation model
The official PyTorch implementation of Google's Gemma models
Inference code for scalable emulation of protein equilibrium ensembles
Audio Language Models are Few-Shot Learners
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Collection of Gemma 3 variants that are trained for performance
High-resolution models for human tasks
Ling is a MoE LLM provided and open-sourced by InclusionAI
Personalize Any Characters with a Scalable Diffusion Transformer
Bidirectional token-classification model for identifiable info
Pretrained time-series foundation model developed by Google Research
A Production-ready Reinforcement Learning AI Agent Library
A PyTorch library for implementing flow matching algorithms
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Controllable & emotion-expressive zero-shot TTS
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion