Language modeling in a sentence representation space
VMZ: Model Zoo for Video Modeling
Qwen3-Coder is the code version of Qwen3
26m function call model that runs on incredibly small devices
Convert Google Gemini web into OpenAI-compatible API
gpt-oss-120b and gpt-oss-20b are two open-weight language models
ICLR2024 Spotlight: curation/training code, metadata, distribution
PyTorch code and models for the DINOv2 self-supervised learning
4M: Massively Multimodal Masked Modeling
Qwen-Image is a powerful image generation foundation model
GLM-4 series: Open Multilingual Multimodal Chat LMs
Claude Code image, a one-stop open source transit service
Reference PyTorch implementation and models for DINOv3
Memory-efficient and performant finetuning of Mistral's models
Open Multilingual Multimodal Chat LMs
Chat & pretrained large audio language model proposed by Alibaba Cloud
Software that can generate photos from paintings
A mix of GAN implementations including progressive growing
Compact 3B-param multimodal model for efficient on-device reasoning
Compact 8B multimodal instruct model optimized for edge deployment
Efficient MoE model for reasoning, coding, and AI agent workflows
Flexible text-to-text transformer model for multilingual NLP tasks
Efficient 8B multimodal model tuned for advanced reasoning tasks.
High-precision 14B multimodal model built for advanced reasoning tasks
Ultra-efficient 3B multimodal instruct model built for edge deployment