Long-form streaming TTS system for multi-speaker dialogue generation
Snippet solution for Vim
The behavior guidance framework for customer-facing LLM agents
Open-source image generative foundation model
SQL-Driven RAG Engine
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Handwritten Text Recognition (HTR) system implemented with TensorFlow
An Open Source text-to-speech system built by inverting Whisper
A modular voice assistant application for experimenting
Tools like web browser, computer access and code runner for LLMs
lightweight package to simplify LLM API calls
An open-source toolkit for monitoring Language Learning Models (LLMs)
Apache-2.0 open-source image generation and editing model family
Powerful Android AI agent with tools, automation, and Linux shell
AI-powered tool for generating, optimizing, and translating subtitles
A community-supported supercharged version of paperless
Qwen-Image is a powerful image generation foundation model
Miso TTS is an 8 billion, highly emotive text-to-speech model
A fast TTS architecture with conditional flow matching
Free, high-quality text-to-speech API endpoint to replace OpenAI
CLIP, Predict the most relevant text snippet given an image
Multimodal-Driven Architecture for Customized Video Generation
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Knowledge Agents and Management in the Cloud
The most accurate natural language detection library for Python