AirLLM 70B inference with single 4GB GPU
Unified KV Cache Compression Methods for Auto-Regressive Models
Neural Network architecture based on ideas of the original LSTM
Real-time NVIDIA GPU dashboard
Open-source large language model family from Tencent Hunyuan
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
Redundancy-aware KV Cache Compression for Reasoning Models
LangChain4j is an open-source Java library
AI Agent Development Guide, LangGraph in Action, Advanced RAG
High-performance inference framework for large language models
Fast and efficient unstructured data extraction
On the Structural Pruning of Large Language Models
DepGraph: Towards Any Structural Pruning
Serving multiple LoRA finetuned LLM as one
A large model training tool that supports training large models
Calculate token/s & GPU memory requirement for any LLM
Flagship MoE model for long-context agents and complex coding