Qwen3-TTS is an open-source series of TTS models
Achieving 3+ generation speedup on reasoning tasks
The most powerful local music generation model
Lets make video diffusion practical
Ultra-Efficient LLMs on End Device
Fast and Universal 3D reconstruction model for versatile tasks
MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation
Industrial-level controllable zero-shot text-to-speech system
Sharp Monocular Metric Depth in Less Than a Second
Image generation model with single-stream diffusion transformer
PyTorch code and models for the DINOv2 self-supervised learning
Hunyuan Translation Model Version 1.5
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
Block Diffusion for Ultra-Fast Speculative Decoding
A Unified Framework for Text-to-3D and Image-to-3D Generation
State of the art LLM and coding model
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
FlashMLA: Efficient Multi-head Latent Attention Kernels
Powerful open source image generation model
Official DeiT repository
GLM-130B: An Open Bilingual Pre-Trained Model (ICLR 2023)
A method to increase the speed and lower the memory footprint