Port of OpenAI's Whisper model in C/C++
Run Local LLMs on Any Device. Open-source
Run frontier MoE models on hardware you already own
Unofficial (Golang) Go bindings for the Hugging Face Inference API
Low-latency REST API for serving text-embeddings
The free, Open Source alternative to OpenAI, Claude and others
Easiest and laziest way for building multi-agent LLMs applications
Optimizing inference proxy for LLMs
User-friendly AI Interface
The Triton Inference Server provides an optimized cloud
Unified Model Serving Framework
Private Open AI on Kubernetes
A library for accelerating Transformer models on NVIDIA GPUs
Operating LLMs in production
Deep Learning API and Server in C++14 support for Caffe, PyTorch
Run local LLMs like llama, deepseek, kokoro etc. inside your browser
Python Package for ML-Based Heterogeneous Treatment Effects Estimation
Bring the notion of Model-as-a-Service to life
Large Language Model Text Generation Inference
Simplifies the local serving of AI models from any source
A RWKV management and startup tool, full automation, only 8MB
Replace OpenAI GPT with another LLM in your app
OpenAI swift async text to image for SwiftUI app using OpenAI
Library for OCR-related tasks powered by Deep Learning
Data manipulation and transformation for audio signal processing