A library for accelerating Transformer models on NVIDIA GPUs
A high-throughput and memory-efficient inference and serving engine
950 line, minimal, extensible LLM inference engine built from scratch
A lightweight vLLM implementation built from scratch
TokenSpeed is a speed-of-light LLM inference engine
Pruna is a model optimization framework built for developers
High-performance inference framework for large language models
LightLLM is a Python-based LLM (Large Language Model) inference
Code for running inference and finetuning with SAM 3 model
Low-latency AI inference engine optimized for mobile devices
Offline inference engine for art, real-time voice conversations
Parallax is a distributed model serving framework
RGBD video generation model conditioned on camera input
Universal LLM Deployment Engine with ML Compilation
Trainable latent-memory framework for 100M-token contexts
Tensor search for humans
Supercharge Your LLM with the Fastest KV Cache Layer
Effortless data labeling with AI support from Segment Anything
Multi-Agent daTa geneRation Infra and eXperimentation framework
Non-autoregressive System 1 decision engine
Enables the best performance on NVIDIA RTX Graphics Cards
Superduper: Integrate AI models and machine learning workflows
Running large language models on a single GPU
Inference Llama 2 in one file of pure C
Toolbox of models, callbacks, and datasets for AI/ML researchers