AI assistant that supports knowledge bases, model APIs
Fast, local-first web content extraction for LLMs
Fast, flexible LLM inference
Query anything (GitHub, Notion, +40 more) with SQL and let LLMs
LLM inference in C/C++
LLM inference server with continuous batching & SSD caching
Run AI models locally on your machine with node.js bindings for llama
High-speed Large Language Model Serving for Local Deployment
Quick illustration of how one can easily read books together with LLMs
Fast and efficient unstructured data extraction
Masks sensitive data and secrets before they reach AI
local-first semantic code search engine
ChatGLM3 series: Open Bilingual Chat LLMs | Open Source Bilingual Chat
AI search engine - self-host with local or cloud LLMs
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
Web app for interacting with any LangGraph agent (PY & TS) via a chat
TokenSpeed is a speed-of-light LLM inference engine
Sovereign, Local-First Agentic AI Linux OS for Toughbook CF-52.
Fully private LLM chatbot that runs entirely with a browser
Auto-GPT on the browser
This website is a free, open-source guide on prompt engineering
Chat with local GGUF LLMs on your own machine