Framework for building real-time voice and multimodal AI agents
A list of free LLM inference resources accessible via API
Offline inference engine for art, real-time voice conversations
Interface for OuteTTS models
A TTS model capable of generating ultra-realistic dialogue
Open-source image generative foundation model
The official Python library for the Fish Audio API
A 0.1B Omni model trained from scratch
Open Source Speech Language Model
Multimodal embedding and reranking models built on Qwen3-VL
NLP Cloud serves high performance pre-trained or custom models for NER
Stable Diffusion built-in to Blender
ProtoMotions is a GPU-accelerated simulation and learning framework
An open source implementation of CLIP
Open source healthcare AI
A high-quality PDF to Markdown tool based on large language model
A full spaCy pipeline and models for scientific/biomedical documents
Sample code and notebooks for Generative AI on Google Cloud
Enhances Tesseract OCR output using LLMs (local or API)
A high-quality rapid TTS voice cloning model
The most powerful local music generation model
Implementation of Phenaki Video, which uses Mask GIT
On-device TTS model by Neuphonic
Large-language-model & vision-language-model based on Linear Attention
A modular voice assistant application for experimenting