AI assistant that supports knowledge bases, model APIs
Fast, local-first web content extraction for LLMs
Fast, flexible LLM inference
Query anything (GitHub, Notion, +40 more) with SQL and let LLMs
LLM inference in C/C++
LLM inference server with continuous batching & SSD caching
Run AI models locally on your machine with node.js bindings for llama
High-speed Large Language Model Serving for Local Deployment
Quick illustration of how one can easily read books together with LLMs
Masks sensitive data and secrets before they reach AI
Fast and efficient unstructured data extraction
local-first semantic code search engine
ChatGLM3 series: Open Bilingual Chat LLMs | Open Source Bilingual Chat
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
AI search engine - self-host with local or cloud LLMs
Web app for interacting with any LangGraph agent (PY & TS) via a chat
TokenSpeed is a speed-of-light LLM inference engine
Fully private LLM chatbot that runs entirely with a browser
Auto-GPT on the browser
This website is a free, open-source guide on prompt engineering