Generate audiobooks from EPUBs, PDFs and text with captions
A robust, efficient, low-latency speech-to-text library
Easily turn large sets of image urls to an image dataset
Simple HTML5, YouTube and Vimeo player
Abstraction layer over YouTube's internal API
Automated YouTube Shorts pipeline
Give Claude the ability to watch any video
Let's use AI to Earn
CLIP, Predict the most relevant text snippet given an image
4M: Massively Multimodal Masked Modeling
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
OpenAI swift async text to image for SwiftUI app using OpenAI
Auto-Subtitle Generator — Free AI Subtitle & Caption Generator
A state-of-the-art open visual language model
Towards Real-World Vision-Language Understanding
An enhanced HTML 5 file input for Bootstrap 5.x/4.x./3.x
Implementation of Dreambooth
Packages with more than 80 components for all delphi versions
An open-source framework for training large multimodal models
The ultimate tool to automate custom telegram message forwarding
Elegant, responsive, flexible and lightweight modal plugin with jQuery
A simple yet powerful JQuery star rating plugin with fractional rating
A lightweight, dependency-free Python library