State-of-the-art (SoTA) text-to-video pre-trained model
Community plugin marketplace for Claude Cowork and Claude Code
An Efficient Agentic Model for Computer Use
Audio foundation model excelling in audio understanding
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
GLM-4-Voice | End-to-End Chinese-English Conversational Model
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Qwen3-omni is a natively end-to-end, omni-modal LLM
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions
Tiny vision language model
The official PyTorch implementation of Google's Gemma models
SOTA on-device LLMs, small yet powerful
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
An Open Real-time Video-Language Interaction System
A 0.1B Omni model trained from scratch
26m function call model that runs on incredibly small devices
Open Source Speech Language Model
Open-source industrial-grade ASR models
Fast-stable-diffusion + DreamBooth
A Pragmatic VLA Foundation Model
Multimodal embedding and reranking models built on Qwen3-VL
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning