Accelerate local LLM inference and finetuning
A high-throughput and memory-efficient inference and serving engine
Advanced LLM-powered brute-force tool combining AI intelligence
Run Local LLMs on Any Device. Open-source
Compress tool outputs, logs, files, and RAG chunks
lightweight package to simplify LLM API calls
Traditional Mandarin LLMs for Taiwan
File Parser optimised for LLM Ingestion with no loss
SDG is a specialized framework
Simple, Pythonic building blocks to evaluate LLM applications
Adding guardrails to large language models
TokenSpeed is a speed-of-light LLM inference engine
Framework to easily create LLM powered bots over any dataset
Play ChatGPT and other LLM with Xiaomi AI Speaker
Operating LLMs in production
An LLM-powered knowledge curation system that researches topics
A security scanner for custom LLM applications
Replace OpenAI GPT with another LLM in your app
Let Claude (or any LLM) actually watch a video
AirLLM 70B inference with single 4GB GPU
LLM inference server with continuous batching & SSD caching
Easy-to-use LLM fine-tuning framework (LLaMA-2, BLOOM, Falcon
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
⚡ Building applications with LLMs through composability ⚡
Find the local LLM that actually runs and performs best