FlashInfer: Kernel Library for LLM Serving
A RWKV management and startup tool, full automation, only 8MB
C++ library for high performance inference on NVIDIA GPUs
OpenMMLab Model Deployment Framework
Self-contained Machine Learning and Natural Language Processing lib
Deep learning inference framework optimized for mobile platforms
Fast and user-friendly runtime for transformer inference