Redundancy-aware KV Cache Compression for Reasoning Models
Block Diffusion for Ultra-Fast Speculative Decoding
FlashMLA: Efficient Multi-head Latent Attention Kernels
A high-throughput and memory-efficient inference and serving engine
Fast LLM speculative inference server for consumer hardware
Achieving 3+ generation speedup on reasoning tasks
A unified library of SOTA model optimization techniques
tiktoken is a fast BPE tokeniser for use with OpenAI's models
The RF and reverse engineering framework for everyone
Fast Multimodal LLM on Mobile Devices
High-performance Inference and Deployment Toolkit for LLMs and VLMs
Document Image Parsing via Heterogeneous Anchor Prompting”
Official repository for LTX-Video
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
SGLang is a fast serving framework for large language models
Build and train a GPT-style language model
Data manipulation and transformation for audio signal processing
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Audio Language Models are Few-Shot Learners
TensorRT LLM provides users with an easy-to-use Python API
Autonomous Agents (LLMs) research papers. Updated Daily
Official implementation of Watermark Anything with Localized Messages
Scalable data pre processing and curation toolkit for LLMs