FlashInfer: Kernel Library for LLM Serving
Port of Facebook's LLaMA model in C/C++
Open standard for machine learning interoperability
The official Python client for the Huggingface Hub
LLM.swift is a simple and readable library
AI interface for tinkerers (Ollama, Haystack RAG, Python)
A high-performance inference system for large language models
Framework which allows you transform your Vector Database
20+ high-performance LLMs with recipes to pretrain, finetune at scale
Optimizing inference proxy for LLMs
Connect home devices into a powerful cluster to accelerate LLM
Pure C++ implementation of several models for real-time chatting
Deep learning optimization library: makes distributed training easy
MII makes low-latency and high-throughput inference possible
AICI: Prompts as (Wasm) Programs
LLMFlows - Simple, Explicit and Transparent LLM Apps
Framework for Accelerating LLM Generation with Multiple Decoding Heads