The official Python client for the Huggingface Hub
Port of Facebook's LLaMA model in C/C++
Open standard for machine learning interoperability
AI interface for tinkerers (Ollama, Haystack RAG, Python)
FlashInfer: Kernel Library for LLM Serving
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
A high-performance inference system for large language models
Framework which allows you transform your Vector Database
LLM.swift is a simple and readable library
Optimizing inference proxy for LLMs
Pure C++ implementation of several models for real-time chatting
Connect home devices into a powerful cluster to accelerate LLM
20+ high-performance LLMs with recipes to pretrain, finetune at scale
The Triton Inference Server provides an optimized cloud
Bring the notion of Model-as-a-Service to life
MII makes low-latency and high-throughput inference possible
Deep learning optimization library: makes distributed training easy
AICI: Prompts as (Wasm) Programs
LLMFlows - Simple, Explicit and Transparent LLM Apps
Framework for Accelerating LLM Generation with Multiple Decoding Heads