Video understanding codebase from FAIR for reproducing video models
Multimodal-Driven Architecture for Customized Video Generation
Tiny vision language model
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
scikit-learn compatible tabular foundation model
Model export recipes, Python primitives, and Swift runtime utilities
Long-form streaming TTS system for multi-speaker dialogue generation
General-purpose image editing model that delivers high-fidelity
PyTorch code and models for the DINOv2 self-supervised learning
A Customizable Image-to-Video Model based on HunyuanVideo
Open-source large language model family from Tencent Hunyuan
Phi-3.5 for Mac: Locally-run Vision and Language Models
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
1B text generation model based on the HRM architecture
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Accurate × Fast × Comprehensive
26m function call model that runs on incredibly small devices
Qwen3-ASR is an open-source series of ASR models
Collection of Gemma 3 variants that are trained for performance
CLIP, Predict the most relevant text snippet given an image
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
Large Multimodal Models for Video Understanding and Editing
RGBD video generation model conditioned on camera input
Netease Youdao's open-source embedding and reranker models