Play ChatGPT and other LLM with Xiaomi AI Speaker
Python bindings for llama.cpp
Easy-to-use LLM fine-tuning framework (LLaMA-2, BLOOM, Falcon
ContextGem: Effortless LLM extraction from documents
A Model Context Protocol (MCP) server implementation
⚡ Building applications with LLMs through composability ⚡
Optimizing inference proxy for LLMs
AirLLM 70B inference with single 4GB GPU
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
A Survey of Large Language Models
950 line, minimal, extensible LLM inference engine built from scratch
Evaluation and Tracking for LLM Experiments
OpenLIT is an open-source LLM Observability tool
Evaluate your LLM's response with Prometheus and GPT4
Easiest and laziest way for building multi-agent LLMs applications
Open-source observability for your LLM application
Harness LLMs with Multi-Agent Programming
Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT method
Lemonade helps users run local LLMs with the highest performance
Open-source large language model family from Tencent Hunyuan
LLM abstractions that aren't obstructions
LLM inference server with continuous batching & SSD caching
Bridging LLM and Recommender System
A powerful tool for automated LLM fuzzing