Generate audiobooks from EPUBs, PDFs and text with captions
A robust, efficient, low-latency speech-to-text library
Easily turn large sets of image urls to an image dataset
Give Claude the ability to watch any video
CLIP, Predict the most relevant text snippet given an image
4M: Massively Multimodal Masked Modeling
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A simple screen parsing tool towards pure vision based GUI agent
A state-of-the-art open visual language model
Towards Real-World Vision-Language Understanding
Implementation of Dreambooth
An open-source framework for training large multimodal models
The ultimate tool to automate custom telegram message forwarding
Official implementation for UniVL video and language training models