AirLLM 70B inference with single 4GB GPU
Unified KV Cache Compression Methods for Auto-Regressive Models
Neural Network architecture based on ideas of the original LSTM
Real-time NVIDIA GPU dashboard
Open-source large language model family from Tencent Hunyuan
Accessible large language models via k-bit quantization for PyTorch
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
Redundancy-aware KV Cache Compression for Reasoning Models
High-speed Large Language Model Serving for Local Deployment
Drag & drop UI to build your customized LLM flow
LLM training in simple, raw C/CUDA
LangChain4j is an open-source Java library
AI Agent Development Guide, LangGraph in Action, Advanced RAG
High-performance inference framework for large language models
Fast and efficient unstructured data extraction
On the Structural Pruning of Large Language Models
DepGraph: Towards Any Structural Pruning
Capable of understanding text, audio, vision, video
Serving multiple LoRA finetuned LLM as one
A large model training tool that supports training large models
Calculate token/s & GPU memory requirement for any LLM
Chat with LLM like Vicuna totally in your browser with WebGPU
Flagship MoE model for long-context agents and complex coding