MII makes low-latency and high-throughput inference possible
An Open Source text-to-speech system built by inverting Whisper
Qwen3-omni is a natively end-to-end, omni-modal LLM
Stable Diffusion built-in to Blender
Easy Docker setup for Stable Diffusion with user-friendly UI
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
Plug-n-play module turning text-to-image models into animation
Marrying Grounding DINO with Segment Anything & Stable Diffusion
Overcoming Data Limitations for High-Quality Video Diffusion Models
Multi-user UI for managing and running Stable Diffusion workflows tool
ComfyUI nodes for LivePortrait
Run GGUF models easily with a UI or API. One File. Zero Install.
A fast TTS architecture with conditional flow matching
Autoregressive Model Beats Diffusion
Implementation of Video Diffusion Models
numerical simulation code for solving transport equations in 1D/2D/3D
Stable Diffusion with Core ML on Apple Silicon
Release for Improved Denoising Diffusion Probabilistic Models
Stable-diffusion-webui-pixelization
Towards Human-Level Text-to-Speech through Style Diffusion
Virtual AI anchor that combines state-of-the-art technology
Removes backgrounds from pictures. Extension for webui
WebUI extension for ControlNet
Implementation of Recurrent Interface Network (RIN)
Fast ODE Solver for Diffusion Probabilistic Model Sampling