1B text generation model based on the HRM architecture
Awesome multilingual OCR toolkits based on PaddlePaddle
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Industrial-level controllable zero-shot text-to-speech system
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Official inference repo for FLUX.1 models
Generating Immersive, Explorable, and Interactive 3D Worlds
Open-source multi-speaker long-form text-to-speech model
State-of-the-art (SoTA) text-to-video pre-trained model
A Family of Open Sourced Music Foundation Models
State-of-the-art TTS model under 25MB
Qwen3-omni is a natively end-to-end, omni-modal LLM
Controllable & emotion-expressive zero-shot TTS
Multimodal-Driven Architecture for Customized Video Generation
Miso TTS is an 8 billion, highly emotive text-to-speech model
Collection of Gemma 3 variants that are trained for performance
AI PPT Track Terminator, the strongest PPT Skill ever
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Large-language-model & vision-language-model based on Linear Attention
The most powerful local music generation model
LLM-based Reinforcement Learning audio edit model
Official Python inference and LoRA trainer package
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Fast stable diffusion on CPU and AI PC
HY-Motion model for 3D character animation generation