Flet enables developers to easily build realtime web and mobile apps
Robust Speech Recognition Across Languages, Dialects
Fast multimodal LLM for real-time voice interaction and AI apps
Controllable & emotion-expressive zero-shot TTS
Math OCR model that outputs LaTeX and markdown
Qwen's most powerful open-source image generation model
Handwritten Text Recognition (HTR) system implemented with TensorFlow
Reading book source
Code for running inference and finetuning with SAM 3 model
Voice Recognition to Text Tool
Open-source image generative foundation model
Search all of YouTube from the command line
Speakr is a personal, self-hosted web application
Miso TTS is an 8 billion, highly emotive text-to-speech model
Snippet solution for Vim
A community-supported supercharged version of paperless
AI-powered code assistant for Vim. OpenAI and ChatGPT plugin for Vim
Capable of understanding text, audio, vision, video
Python bindings for MuPDF's rendering library.
State-of-the-art (SoTA) text-to-video pre-trained model
1 min voice data can also be used to train a good TTS model
RAG-Anything: All-in-One RAG Framework
The simplest, fastest repository for training/finetuning models
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Context-aware desktop AI assistant that understands screen content