Run Bonsai (1-bit) and Ternary-Bonsai language models locally
InvokeAI is a leading creative engine for Stable Diffusion models
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Focus on prompting and generating
Gemma open-weight LLM library, from Google DeepMind
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Spark-TTS Inference Code
High-Resolution Image Synthesis with Latent Diffusion Models
The Multi-Agent Framework
TTS model capable of streaming conversational audio in realtime
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
AI generative media user experience highlighting use of APIs
Miso TTS is an 8 billion, highly emotive text-to-speech model
Accurate × Fast × Comprehensive
Recovering the Visual Space from Any Views
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
State-of-the-art diffusion models for image and audio generation
RGBD video generation model conditioned on camera input
Dealing with all unstructured data, such as reverse image search
Multilingual sentence & image embeddings with BERT
A high-quality PDF to Markdown tool based on large language model
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
Open-source framework for intelligent speech interaction
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Temporal-Consistent Diffusion Model for Real-World Video