Port of Facebook's LLaMA model in C/C++
Port of OpenAI's Whisper model in C/C++
A high-throughput and memory-efficient inference and serving engine
Run Local LLMs on Any Device. Open-source
ONNX Runtime: cross-platform, high performance ML inferencing
High-performance neural network inference framework for mobile
OpenVINO™ Toolkit repository
The free, Open Source alternative to OpenAI, Claude and others
LLMs as Copilots for Theorem Proving in Lean
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
User-friendly AI Interface
Everything you need to build state-of-the-art foundation models
Run serverless GPU workloads with fast cold starts on bare-metal
Sparsity-aware deep learning inference runtime for CPUs
State-of-the-art diffusion models for image and audio generation
Open-Source AI Camera. Empower any camera/CCTV
Pure C++ implementation of several models for real-time chatting
An Open-Source Programming Framework for Agentic AI
A unified framework for scalable computing
Fast inference engine for Transformer models
A lightweight vision library for performing large object detection
Libraries for applying sparsification recipes to neural networks
Unofficial (Golang) Go bindings for the Hugging Face Inference API
Protect and discover secrets using Gitleaks
A GPU-accelerated library containing highly optimized building blocks