FlashInfer: Kernel Library for LLM Serving
C++ library for high performance inference on NVIDIA GPUs
A high-performance inference system for large language models
State-of-the-art Parameter-Efficient Fine-Tuning
ONNX Runtime: cross-platform, high performance ML inferencing
Large Language Model Text Generation Inference
Port of Facebook's LLaMA model in C/C++
20+ high-performance LLMs with recipes to pretrain, finetune at scale
Uncover insights, surface problems, monitor, and fine tune your LLM
Run Local LLMs on Any Device. Open-source
Run frontier MoE models on hardware you already own
OpenVINO™ Toolkit repository
High-performance neural network inference framework for mobile
lightweight, standalone C++ inference engine for Google's Gemma models
Run serverless GPU workloads with fast cold starts on bare-metal
Fast inference engine for Transformer models
An Open-Source Programming Framework for Agentic AI
Run local LLMs like llama, deepseek, kokoro etc. inside your browser
Optimizing inference proxy for LLMs
On-device AI across mobile, embedded and edge for PyTorch
Connect home devices into a powerful cluster to accelerate LLM
Unified Model Serving Framework
A high-performance ML model serving framework, offers dynamic batching
A RWKV management and startup tool, full automation, only 8MB
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference