Unified Multimodal Understanding and Generation Models
FAIR Sequence Modeling Toolkit 2
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Renderer for the harmony response format to be used with gpt-oss
A series of math-specific large language models of our Qwen2 series
Repo of Qwen2-Audio chat & pretrained large audio language model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
The Clay Foundation Model - An open source AI model and interface
Netease Youdao's open-source embedding and reranker models
An Efficient Agentic Model for Computer Use
Phi-3.5 for Mac: Locally-run Vision and Language Models
Infinite Worlds with Versatile Interactions
scikit-learn compatible tabular foundation model
1B text generation model based on the HRM architecture
Robust Speech Recognition Across Languages, Dialects
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Visual Causal Flow
The official PyTorch implementation of Google's Gemma models
Programmatic access to the AlphaGenome model
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
An Open Real-time Video-Language Interaction System