Powerful Android AI agent with tools, automation, and Linux shell
Qwen3-omni is a natively end-to-end, omni-modal LLM
A lightweight text-to-speech model with zero-shot voice cloning
Open-source multi-speaker long-form text-to-speech model
Style-Bert-VITS2: Bert-VITS2 with more controllable voice styles
Label Studio is a multi-type data labeling and annotation tool
AI PPT Track Terminator, the strongest PPT Skill ever
CLIP, Predict the most relevant text snippet given an image
AI-powered tool for generating, optimizing, and translating subtitles
Unifying 3D Mesh Generation with Language Models
Multimodal-Driven Architecture for Customized Video Generation
StreamSpeech is a seamless model for offline speech recognition
Official MiniMax Model Context Protocol (MCP) server
A community-supported supercharged version of paperless
1 min voice data can also be used to train a good TTS model
Apache-2.0 open-source image generation and editing model family
Python library and CLI tool to interface with Google Translate
Implementation of Imagen, Google's Text-to-Image Neural Network
Python binding to the Apache Tika™ REST services
World's first open-source, agentic video production system
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Speech recognition module for Python
Controllable & emotion-expressive zero-shot TTS
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
A robust, efficient, low-latency speech-to-text library