OCR expert VLM powered by Hunyuan's native multimodal architecture
Qwen-Image is a powerful image generation foundation model
Video understanding codebase from FAIR for reproducing video models
PyTorch code and models for the DINOv2 self-supervised learning
code for Mesh R-CNN, ICCV 2019
Foundation Models for Time Series
Official implementation of Watermark Anything with Localized Messages
Bidirectional token-classification model for identifiable info
A Multi-Modal World Model for Reconstructing, Generating, Simulation
The ChatGPT Retrieval Plugin lets you easily find personal documents
Dataset of GPT-2 outputs for research in detection, biases, and more
This repository contains the official implementation of research
PyTorch implementation of MAE
PyTorch implementation of YOLOv4