A high-throughput and memory-efficient inference and serving engine
Deep learning optimization library: makes distributed training easy
Parallax is a distributed model serving framework
Minimal Python framework for scalable AI inference servers fast
High-performance inference server for text embeddings models API layer
Running large language models on a single GPU
AI memory OS for LLM and Agent systems
Low-latency REST API for serving text-embeddings
MII makes low-latency and high-throughput inference possible
Real-Time Open-Ended Video Editing with Autoregressive Diffusion
A new kind of Progress Bar, with real-time throughput, ETA
950 line, minimal, extensible LLM inference engine built from scratch
Large Language Model Text Generation Inference
IPTV live stream source automatic update tool
The async Python driver for MongoDB and Tornado or asyncio
CoreNet: A library for training deep neural networks
Fast and memory-efficient exact attention
Supercharge Your LLM with the Fastest KV Cache Layer
DeepEP: an efficient expert-parallel communication library
Lets make video diffusion practical
Windows multi-NIC bandwidth aggregator
Lemonade helps users run local LLMs with the highest performance
TensorRT LLM provides users with an easy-to-use Python API
slime is an LLM post-training framework for RL Scaling
Scrape tweets, profiles, followers and following from Twitter/X