State-of-the-art TTS model under 25MB
Powerful AI language model (MoE) optimized for efficiency/performance
Open-source, high-performance AI model with advanced reasoning
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
A trainable PyTorch reproduction of AlphaFold 3
Open-source multi-speaker long-form text-to-speech model
Awesome multilingual OCR toolkits based on PaddlePaddle
Python inference and LoRA trainer package for the LTX-2 audio–video
Official Python inference and LoRA trainer package
Robust Speech Recognition Across Languages, Dialects
Open Source Speech Language Model
Z80-μLM is a 2-bit quantized language model
Text and image to video generation: CogVideoX and CogVideo
Qwen3-TTS is an open-source series of TTS models
Models for object and human mesh reconstruction
Accurate × Fast × Comprehensive
High-Resolution Image Synthesis with Latent Diffusion Models
AI cognitive-enhancement Skills based on Anthropic's J-space
A Powerful Native Multimodal Model for Image Generation
Netease Youdao's open-source embedding and reranker models
Video understanding codebase from FAIR for reproducing video models
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Unified Multimodal Understanding and Generation Models
Audio Language Models are Few-Shot Learners
Open-source framework for intelligent speech interaction