Generate audiobooks from EPUBs, PDFs and text with captions
A robust, efficient, low-latency speech-to-text library
Easily turn large sets of image urls to an image dataset
Automated YouTube Shorts pipeline
Give Claude the ability to watch any video
CLIP, Predict the most relevant text snippet given an image
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
4M: Massively Multimodal Masked Modeling
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
A simple screen parsing tool towards pure vision based GUI agent
A state-of-the-art open visual language model
Towards Real-World Vision-Language Understanding
Implementation of Dreambooth
An open-source framework for training large multimodal models
The ultimate tool to automate custom telegram message forwarding
A lightweight, dependency-free Python library
Official implementation for UniVL video and language training models
Easily and Quickly add Captions to your photos
FUSE-based filesystem reflecting XWindows into files
Browser interface to your memories