Browse free open source Python LLM Inference Tools and projects below. Use the toggles on the left to filter open source Python LLM Inference Tools by OS, license, language, programming language, and project status.
Run Local LLMs on Any Device. Open-source
A high-throughput and memory-efficient inference and serving engine
OpenMMLab Model Deployment Framework
FlashInfer: Kernel Library for LLM Serving
Lightweight anchor-free object detection model
Ready-to-use OCR with 80+ supported languages
Everything you need to build state-of-the-art foundation models
DoWhy is a Python library for causal inference
MII makes low-latency and high-throughput inference possible
Low-latency REST API for serving text-embeddings
Bring the notion of Model-as-a-Service to life
Multi-Modal Neural Networks for Semantic Search, based on Mid-Fusion
PyTorch library of curated Transformer models and their components
Open platform for training, serving, and evaluating language models
20+ high-performance LLMs with recipes to pretrain, finetune at scale
A high-performance ML model serving framework, offers dynamic batching
Lightweight Python library for adding real-time multi-object tracking
A toolkit to optimize ML models for deployment for Keras & TensorFlow
The Triton Inference Server provides an optimized cloud
The official Python client for the Huggingface Hub
Deep learning optimization library: makes distributed training easy
Library for OCR-related tasks powered by Deep Learning
OpenMMLab Video Perception Toolbox
PyTorch extensions for fast R&D prototyping and Kaggle farming
A unified framework for scalable computing