Instant voice cloning by MIT and MyShell. Audio foundation model
Automatic subtitle synchronization tool
A Family of Open Sourced Music Foundation Models
Zero-copy PDF text extraction library written in Zig
A tool for semi-automatic cell type classification, harmonization
SOTA Open Source TTS
Interface for OuteTTS models
Calculate quality metrics with FFmpeg (SSIM, PSNR, VMAF, VIF)
A lightweight text-to-speech model with zero-shot voice cloning
An open-source toolkit for BigMac-style pipeline-parallel training
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Taming Stable Diffusion for Lip Sync
Run PyTorch LLMs locally on servers, desktop and mobile
TorchMultimodal is a PyTorch library
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
A collection of tools, libraries, and tests for Vulkan shader
Multi-lingual large voice generation model, providing inference
Learn all about Digital Forensics and Computer Forensics
Official code for Style Aligned Image Generation via Shared Attention
Open source implementation of Microsoft's VALL-E X zero-shot TTS model
Sample code for Google Cloud Vision
Clone a voice in 5 seconds to generate arbitrary speech in real-time
The basic distribution probability Tutorial for Deep Learning Research