Infinite Worlds with Versatile Interactions
scikit-learn compatible tabular foundation model
1B text generation model based on the HRM architecture
The official PyTorch implementation of Google's Gemma models
Programmatic access to the AlphaGenome model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
4M: Massively Multimodal Masked Modeling
Hackable and optimized Transformers building blocks
Official implementation of DreamCraft3D
Qwen2.5-VL is the multimodal large language model series
New family of code large language models (LLMs)
Controllable & emotion-expressive zero-shot TTS
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Renderer for the harmony response format to be used with gpt-oss
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
Audio Language Models are Few-Shot Learners
Open-source industrial-grade ASR models
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Hunyuan Translation Model Version 1.5
Multimodal embedding and reranking models built on Qwen3-VL
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling