Flexible Photo Recrafting While Preserving Your Identity
Multimodal-Driven Architecture for Customized Video Generation
Open-source, code-first Python toolkit for building, evaluating, etc.
Qwen-Image is a powerful image generation foundation model
Instant voice cloning by MIT and MyShell. Audio foundation model
Official inference repo for FLUX.2 models
A Unified Framework for Image Customization
Bindu: Turn any AI agent into a living microservice
One-stop solution for creating your digital avatar from chat history
Personalize Any Characters with a Scalable Diffusion Transformer
A Customizable Image-to-Video Model based on HunyuanVideo
The common language for platforms, agents and businesses.
A Universal Customization Method for Single and Multi Conditioning
Pushing the Frontier of Long Audio-Visual Generation
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Asynchronous coordination layer for AI coding agents
AI agent microservice
Interface for OuteTTS models
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
MARS5 speech model (TTS) from CAMB.AI
Unofficial implementation of InstantID for ComfyUI
Implementation of Make-A-Video, new SOTA text to video generator
FaceOnLive Open KYC: Streamlining Identity Verification with AI
CoTracker is a model for tracking any point (pixel) on a video
Towards Robust Blind Face Restoration with Codebook Lookup Transformer