PyTorch code and models for the DINOv2 self-supervised learning
Recovering the Visual Space from Any Views
Open-Source Financial Large Language Models
Video Object and Interaction Deletion
VMZ: Model Zoo for Video Modeling
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
GLM-4 series: Open Multilingual Multimodal Chat LMs
Audio foundation model excelling in audio understanding
Revolutionizing Database Interactions with Private LLM Technology
Programmatic access to the AlphaGenome model
Unified Multimodal Understanding and Generation Models
Open-source industrial-grade ASR models
Foundation model for image generation
A Pragmatic VLA Foundation Model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Hunyuan Translation Model Version 1.5
Multimodal embedding and reranking models built on Qwen3-VL
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
A Unified Framework for Text-to-3D and Image-to-3D Generation
Personalize Any Characters with a Scalable Diffusion Transformer
OCR expert VLM powered by Hunyuan's native multimodal architecture
Bidirectional token-classification model for identifiable info
Genome modeling and design across all domains of life