Infinite Worlds with Versatile Interactions
scikit-learn compatible tabular foundation model
1B text generation model based on the HRM architecture
Robust Speech Recognition Across Languages, Dialects
The official PyTorch implementation of Google's Gemma models
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
An Open Real-time Video-Language Interaction System
A 0.1B Omni model trained from scratch
Open Source Speech Language Model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Hunyuan Translation Model Version 1.5
Multimodal embedding and reranking models built on Qwen3-VL
Collection of Gemma 3 variants that are trained for performance
Implementation of "MobileCLIP" CVPR 2024
Official implementation of Watermark Anything with Localized Messages
High-resolution models for human tasks
Video understanding codebase from FAIR for reproducing video models
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
Multimodal Diffusion with Representation Alignment
Personalize Any Characters with a Scalable Diffusion Transformer
MOSS‑TTS Family open‑source speech and sound generation model
Bidirectional token-classification model for identifiable info