Tensor search for humans
A high-throughput and memory-efficient inference and serving engine
A lightweight vLLM implementation built from scratch
Ongoing research training transformer models at scale
CV, NLP, LLM project applications, and advanced engineering deployment
A course of learning LLM inference serving on Apple Silicon
Large-language-model & vision-language-model based on Linear Attention