AirLLM 70B inference with single 4GB GPU
Unified KV Cache Compression Methods for Auto-Regressive Models
Neural Network architecture based on ideas of the original LSTM
Open-source large language model family from Tencent Hunyuan
Accessible large language models via k-bit quantization for PyTorch
Redundancy-aware KV Cache Compression for Reasoning Models
LLM training in simple, raw C/CUDA
The official repo of Qwen chat & pretrained large language model
High-performance inference framework for large language models
On the Structural Pruning of Large Language Models
DepGraph: Towards Any Structural Pruning
Capable of understanding text, audio, vision, video
Serving multiple LoRA finetuned LLM as one
A large model training tool that supports training large models