Browse free open source LLM Inference tools and projects for Linux below. Use the toggles on the left to filter open source LLM Inference tools by OS, license, language, programming language, and project status.
Port of OpenAI's Whisper model in C/C++
User-friendly AI Interface
Port of Facebook's LLaMA model in C/C++
Run Local LLMs on Any Device. Open-source
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine
ONNX Runtime: cross-platform, high performance ML inferencing
The free, Open Source alternative to OpenAI, Claude and others
High-performance neural network inference framework for mobile
Open-Source AI Camera. Empower any camera/CCTV
Pure C++ implementation of several models for real-time chatting
A library for accelerating Transformer models on NVIDIA GPUs
AIMET is a library that provides advanced quantization and compression
OpenVINO™ Toolkit repository
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
Protect and discover secrets using Gitleaks
Fast inference engine for Transformer models
On-device AI across mobile, embedded and edge for PyTorch
A RWKV management and startup tool, full automation, only 8MB
Deep learning optimization library: makes distributed training easy
LLMs as Copilots for Theorem Proving in Lean
Official inference library for Mistral models
A scalable inference server for models optimized with OpenVINO
AICI: Prompts as (Wasm) Programs
Run local LLMs like llama, deepseek, kokoro etc. inside your browser