1B text generation model based on the HRM architecture
High-performance inference server for text embeddings models API layer
Document (PDF, Word, PPTX ...) extraction and parse API
The official Python library for the Fish Audio API
Hypernetworks that adapt LLMs for specific benchmark tasks
TTS with kokoro and onnx runtime
Python library and CLI tool to interface with Google Translate
Use Microsoft Edge's online text-to-speech service from Python
MiniMax H3 is a general-purpose, omni-modal generative system
High-Quality Voice Cloning TTS for 600+ Languages
Official inference repo for FLUX.1 models
A robust, efficient, low-latency speech-to-text library
A TTS that fits in your CPU (and pocket)
Offline inference engine for art, real-time voice conversations
A nearly-live implementation of OpenAI's Whisper
Robust Speech Recognition via Large-Scale Weak Supervision
Code for running inference and finetuning with SAM 3 model
Qwen3-TTS is an open-source series of TTS models
A lightweight text-to-speech model with zero-shot voice cloning
Contexts Optical Compression
Google Gen AI Python SDK provides an interface for developers
Industrial-level controllable zero-shot text-to-speech system
The python library for real-time communication
A Family of Open Sourced Music Foundation Models
Open-source multi-speaker long-form text-to-speech model