AI PPT Track Terminator, the strongest PPT Skill ever
An Open Real-time Video-Language Interaction System
Audio foundation model excelling in audio understanding
Video understanding codebase from FAIR for reproducing video models
CLIP, Predict the most relevant text snippet given an image
scikit-learn compatible tabular foundation model
Open Source Speech Language Model
Multimodal embedding and reranking models built on Qwen3-VL
A Production-ready Reinforcement Learning AI Agent Library
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
A SOTA open-source image editing model