Open-source image generative foundation model
The behavior guidance framework for customer-facing LLM agents
SQL-Driven RAG Engine
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Handwritten Text Recognition (HTR) system implemented with TensorFlow
An Open Source text-to-speech system built by inverting Whisper
The next generation Javascript WYSIWYG HTML Editor
A modular voice assistant application for experimenting
The python library for real-time communication
Tools like web browser, computer access and code runner for LLMs
AI-powered tool for generating, optimizing, and translating subtitles
A community-supported supercharged version of paperless
lightweight package to simplify LLM API calls
An open-source toolkit for monitoring Language Learning Models (LLMs)
Apache-2.0 open-source image generation and editing model family
Powerful Android AI agent with tools, automation, and Linux shell
Qwen-Image is a powerful image generation foundation model
Miso TTS is an 8 billion, highly emotive text-to-speech model
Free, high-quality text-to-speech API endpoint to replace OpenAI
A fast TTS architecture with conditional flow matching
CLIP, Predict the most relevant text snippet given an image
Multimodal-Driven Architecture for Customized Video Generation
Dicio assistant app for Android
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Knowledge Agents and Management in the Cloud