A library for accelerating Transformer models on NVIDIA GPUs
A high-throughput and memory-efficient inference and serving engine
A lightweight vLLM implementation built from scratch
950 line, minimal, extensible LLM inference engine built from scratch
High-performance inference framework for large language models
TokenSpeed is a speed-of-light LLM inference engine
Code for running inference and finetuning with SAM 3 model
RGBD video generation model conditioned on camera input
Low-latency AI inference engine optimized for mobile devices
Pruna is a model optimization framework built for developers
LightLLM is a Python-based LLM (Large Language Model) inference
Parallax is a distributed model serving framework
Universal LLM Deployment Engine with ML Compilation
Offline inference engine for art, real-time voice conversations
Tensor search for humans
Enables the best performance on NVIDIA RTX Graphics Cards
Multi-Agent daTa geneRation Infra and eXperimentation framework
Effortless data labeling with AI support from Segment Anything
Running large language models on a single GPU
Superduper: Integrate AI models and machine learning workflows
Supercharge Your LLM with the Fastest KV Cache Layer
Inference Llama 2 in one file of pure C
Toolbox of models, callbacks, and datasets for AI/ML researchers
A Strong and Easy-to-use Single View 3D Hand+Body Pose Estimator
Real-Time State-of-the-art Speech Synthesis for Tensorflow 2