AI assistant that supports knowledge bases, model APIs
LLM inference server with continuous batching & SSD caching
Fast, local-first web content extraction for LLMs
Query anything (GitHub, Notion, +40 more) with SQL and let LLMs
High-speed Large Language Model Serving for Local Deployment
Fast and efficient unstructured data extraction
Quick illustration of how one can easily read books together with LLMs
LLM inference in C/C++
ChatGLM3 series: Open Bilingual Chat LLMs | Open Source Bilingual Chat
Fast, flexible LLM inference
Web app for interacting with any LangGraph agent (PY & TS) via a chat
Masks sensitive data and secrets before they reach AI
AI search engine - self-host with local or cloud LLMs
Run AI models locally on your machine with node.js bindings for llama
local-first semantic code search engine
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
TokenSpeed is a speed-of-light LLM inference engine
Sovereign, Local-First Agentic AI Linux OS for Toughbook CF-52.
Fully private LLM chatbot that runs entirely with a browser
Auto-GPT on the browser
This website is a free, open-source guide on prompt engineering
Chat with local GGUF LLMs on your own machine