GLM-4 series: Open Multilingual Multimodal Chat LMs
Qwen3 is the large language model series developed by Qwen team
Accurate × Fast × Comprehensive
Audio foundation model excelling in audio understanding
Unified Multimodal Understanding and Generation Models
FAIR Sequence Modeling Toolkit 2
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
PyTorch code and models for the DINOv2 self-supervised learning
State-of-the-art TTS model under 25MB
Models for object and human mesh reconstruction
Sharp Monocular Metric Depth in Less Than a Second
An Open Real-time Video-Language Interaction System
Qwen3-ASR is an open-source series of ASR models
A Pragmatic VLA Foundation Model
Python SDK for Claude Agent
Large Multimodal Models for Video Understanding and Editing
RGBD video generation model conditioned on camera input
Code for running inference with the SAM 3D Body Model 3DB
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Generate Any 3D Scene in Seconds
A Production-ready Reinforcement Learning AI Agent Library
Convert Google Gemini web into OpenAI-compatible API
Miso TTS is an 8 billion, highly emotive text-to-speech model
Foundation model for image generation