LM Studio Apple MLX engine
TTS with kokoro and onnx runtime
Fastest LLM inference runtime for Apple Silicon
MCP server for interfacing with Godot game engine
Open source solution that can meet the requirements of workloads
Offline Text To Speech synthesis for python
DeepSeek 4 Flash local inference engine for Metal
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
AI agent harness for AI coding agents
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
SQL-Driven RAG Engine
Universal LLM Deployment Engine with ML Compilation
950 line, minimal, extensible LLM inference engine built from scratch
A self-hostable CDN for databases
Multi-Agent daTa geneRation Infra and eXperimentation framework
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
Run the full 2.78-trillion-parameter Kimi K3 model
Fast Multimodal LLM on Mobile Devices
LiteRT, successor to TensorFlow Lite
Emscripten: An LLVM-to-WebAssembly Compiler
On-device wake word detection powered by deep learning
Smart LLM router
NestJS Helper + AI Chatbot Development
Fast inference engine for Transformer models