Redundancy-aware KV Cache Compression for Reasoning Models
A high-throughput and memory-efficient inference and serving engine
High-performance Inference and Deployment Toolkit for LLMs and VLMs
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Scalable data pre processing and curation toolkit for LLMs
Chinese Llama-3 LLMs) developed from Meta Llama 3
Framework that is dedicated to making neural data processing
Chinese LLaMA & Alpaca large language model + local CPU/GPU training