157 models, 30 providers, one command to find what runs on hardware
High-speed Large Language Model Serving for Local Deployment
LightLLM is a Python-based LLM (Large Language Model) inference
TokenSpeed is a speed-of-light LLM inference engine
Mooncake is the serving platform for Kimi
Apple Intelligence from the command line
Multilingual sentence & image embeddings with BERT
Your Second Brain supercharged by Generative AI
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
How to optimize some algorithm in cuda
A guidance language for controlling large language models
Real-time NVIDIA GPU dashboard
Dance with Intelligence in Your Code
Personal AI Notebooks. Organize files & webpages and generate notes
A New Axis of Sparsity for Large Language Models
A high-performance ML model serving framework, offers dynamic batching
State of the art LLM and coding model
An Easy-to-Use and High-Performance AI Deployment Framework
MobileLLM Optimizing Sub-billion Parameter Language Models
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
ChatGLM2-6B: An Open Bilingual Chat LLM
An efficient forwarding service designed for LLMs
Open-source, developer-first LLMOps platform
Calculate token/s & GPU memory requirement for any LLM
GLM-130B: An Open Bilingual Pre-Trained Model (ICLR 2023)