A simple native web interface that uses ChatTTS to synthesize text
Build multimodal language agents for fast prototype and production
Document Image Parsing via Heterogeneous Anchor Prompting”
Large Multimodal Models for Video Understanding and Editing
StreamSpeech is a seamless model for offline speech recognition
Private AI platform for agents, enterprise search and RAG pipelines
Videomass is a free, open source and cross-platform GUI for FFmpeg
Open-Source Low-Latency Accelerated Linux WebRTC HTML5 Remote Desktop
Spring AI Alibaba examples for building and testing AI apps
WhatsApp MCP server enabling AI access to chats and messaging
A Systematic Framework for Interactive World Modeling
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
State-of-the-art diffusion models for image and audio generation
A high quality MP3 encoder
Build Vision Agents quickly with any model or video provider
Multi-Modal Neural Networks for Semantic Search, based on Mid-Fusion
The data structure for multimodal data
Build AI-powered semantic search applications
pyglet is a cross-platform windowing and multimedia library for Python
Give Claude the ability to watch any video
An Open Source text-to-speech system built by inverting Whisper
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Have a natural, spoken conversation with AI
Strip multi-vendor AI provenance marks
spotify music downloader telegram bot (tracks, albums, playlists)