Block Diffusion for Ultra-Fast Speculative Decoding
FlashMLA: Efficient Multi-head Latent Attention Kernels
Achieving 3+ generation speedup on reasoning tasks
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Official repository for LTX-Video
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Official implementation of Watermark Anything with Localized Messages
Ultra-Efficient LLMs on End Device
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
Facebook AI Research Sequence-to-Sequence Toolkit
Large-scale autoregressive pixel model for image generation by OpenAI
Speculative-decoding accelerator for the 675B Mistral Large 3
High-performance MoE model with MLA, MTP, and multilingual reasoning
Efficient MoE model for reasoning, coding, and AI agent workflows