State-of-the-art TTS model under 25MB
Powerful AI language model (MoE) optimized for efficiency/performance
Open-source, high-performance AI model with advanced reasoning
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Open-source multi-speaker long-form text-to-speech model
Awesome multilingual OCR toolkits based on PaddlePaddle
Python inference and LoRA trainer package for the LTX-2 audio–video
Official Python inference and LoRA trainer package
Robust Speech Recognition Across Languages, Dialects
Open Source Speech Language Model
Z80-μLM is a 2-bit quantized language model
Text and image to video generation: CogVideoX and CogVideo
Qwen3-VL, the multimodal large language model series by Alibaba Cloud
Qwen3-TTS is an open-source series of TTS models
Flux 2 image generation model pure C inference
Models for object and human mesh reconstruction
Accurate × Fast × Comprehensive
AI cognitive-enhancement Skills based on Anthropic's J-space
A Powerful Native Multimodal Model for Image Generation
Netease Youdao's open-source embedding and reranker models
Video understanding codebase from FAIR for reproducing video models
Clean and efficient FP8 GEMM kernels with fine-grained scaling
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Unified Multimodal Understanding and Generation Models
MiniMax M2.1, a SOTA model for real-world dev & agents.