A unified library of SOTA model optimization techniques
AirLLM 70B inference with single 4GB GPU
A modern model graph visualizer and debugger
How to optimize some algorithm in cuda
A modern Anki custom scheduling based on Free Spaced Repetition
Gradient boosting framework based on decision tree algorithms
Designed for training LLM/VLM agents via RL
C++ library for high performance inference on NVIDIA GPUs
Facebook AI Research Sequence-to-Sequence Toolkit written in Python
PyTorch implementation of SimSiam
NVFP4 DiffusionGemma model for fast multimodal text generation