LLM inference server with continuous batching & SSD caching
Quick illustration of how one can easily read books together with LLMs
local-first semantic code search engine
ChatGLM3 series: Open Bilingual Chat LLMs | Open Source Bilingual Chat
TokenSpeed is a speed-of-light LLM inference engine
Chat with local GGUF LLMs on your own machine