Buzz transcribes and translates audio offline
Open-source, high-performance AI model with advanced reasoning
From Images to High-Fidelity 3D Assets
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Official inference repo for FLUX.1 models
The most powerful local music generation model
Awesome multilingual OCR toolkits based on PaddlePaddle
AlphaFold 3 inference pipeline
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Native and Compact Structured Latents for 3D Generation
Official Python inference and LoRA trainer package
AI PPT Track Terminator, the strongest PPT Skill ever
Fast stable diffusion on CPU and AI PC
Industrial-level controllable zero-shot text-to-speech system
Advanced language and coding AI model
An experimental version of DeepSeek model
A SOTA open-source image editing model
Video Object and Interaction Deletion
State-of-the-art TTS model under 25MB
Agentic, Reasoning, and Coding (ARC) foundation models
LTX-Video Support for ComfyUI
Miso TTS is an 8 billion, highly emotive text-to-speech model
A Family of Open Sourced Music Foundation Models
Open-source multi-speaker long-form text-to-speech model
State-of-the-art (SoTA) text-to-video pre-trained model