Give Claude the ability to watch and understand videos
Open-source multi-speaker long-form text-to-speech model
ESP32 Camera motion capture application to record JPEGs to SD card
Qwen2.5-VL is the multimodal large language model series
UME is an in-app debug kits platform for Flutter
Cross Audio-Visual Recognition using 3D Architectures
A modern approach for Computer Vision on the web
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video