TT-NN operator library, and TT-Metalium low level kernel programming
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
A high-performance inference engine for AI models
A.S.E (AICGSecEval) is a repository-level AI-generated code security
A course of learning LLM inference serving on Apple Silicon
The official implementation of RAPTOR
TokenSpeed is a speed-of-light LLM inference engine
Advanced LLM-powered brute-force tool combining AI intelligence
High-performance inference framework for large language models
Weaving the Digital Agent Galaxy
AI-Powered Data Processing: Use LOTUS to process all of your datasets
Mooncake is the serving platform for Kimi
How to optimize some algorithm in cuda
A guidance language for controlling large language models
Advanced language and coding AI model
Build a modern LLM from scratch. Every line commented
Semi-Structured Agentic Framework. Workflows build themselves
Open-source enterprise-level AI knowledge base and MCP
Unified framework for building enterprise RAG pipelines
Scalable data pre processing and curation toolkit for LLMs
csghub-server is the backend server for CSGHub
An extensible framework for Personal Data Management
State of the art LLM and coding model
This repository provides an advanced RAG
Qwen2.5-Coder is the code version of Qwen2.5, the large language model