A high-throughput and memory-efficient inference and serving engine
Deep learning optimization library: makes distributed training easy
Parallax is a distributed model serving framework
High-performance inference server for text embeddings models API layer
Minimal Python framework for scalable AI inference servers fast
Running large language models on a single GPU
AI memory OS for LLM and Agent systems
MII makes low-latency and high-throughput inference possible
Low-latency REST API for serving text-embeddings
950 line, minimal, extensible LLM inference engine built from scratch
Large Language Model Text Generation Inference
Fast and memory-efficient exact attention
Supercharge Your LLM with the Fastest KV Cache Layer
Lemonade helps users run local LLMs with the highest performance
Lets make video diffusion practical
TensorRT LLM provides users with an easy-to-use Python API
slime is an LLM post-training framework for RL Scaling
Accurate × Fast × Comprehensive
Multi-agent autonomous startup system for Claude Code
A TTS that fits in your CPU (and pocket)
High-performance Inference and Deployment Toolkit for LLMs and VLMs
Hackable and optimized Transformers building blocks
Block Diffusion for Ultra-Fast Speculative Decoding
FAIR Sequence Modeling Toolkit 2
Agent framework and applications built upon Qwen>=3.0