High-speed Large Language Model Serving for Local Deployment
Find the local LLM that actually runs and performs best
LLM inference in C/C++
157 models, 30 providers, one command to find what runs on hardware
A high-performance ML model serving framework, offers dynamic batching
Run AI models locally on your machine with node.js bindings for llama
Real-time NVIDIA GPU dashboard
Generate music based on natural language prompts using LLMs
Chinese LLaMA-2 & Alpaca-2 Large Model Phase II Project
Calculate token/s & GPU memory requirement for any LLM
Chinese LLaMA & Alpaca large language model + local CPU/GPU training