scikit-learn compatible tabular foundation model
Lets make video diffusion practical
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Language modeling in a sentence representation space
Official DeiT repository
Learning to Act by Watching Unlabeled Online Videos
PyTorch implementation of MAE
Code for reproducing key results in the paper