Solve puzzles. Learn CUDA
A bidirectional pipeline parallelism algorithm
AirLLM 70B inference with single 4GB GPU
Variational Quantum Circuit Simulator for Quantum Computation Research
A language for fast, portable data-parallel computation
Running large language models on a single GPU
Distributed parallelization of stencil-based GPU and CPU applications
Running a big model on a small laptop
SwissGL is a minimalistic wrapper on top of WebGL2 JS API
Open source machine learning framework
Multi-platform high-performance compute language extension for Rust
Tensor Learning in Python
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
The Zoo Design Studio app
Faster Whisper transcription with CTranslate2
Fast inference engine for Transformer models
Meridian is an MMM framework
Easily compute clip embeddings and build a clip retrieval system
Analyze computation-communication overlap in V3/R1
A high-performance inference engine for AI models
Advanced evolutionary computation library built on top of PyTorch
NeurIPS2025 Spotlight] Quantized Attention
A massively parallel, high-level programming language
Pythonic tool for running machine-learning/high performance workflows
Prevent PyTorch's `CUDA error: out of memory` in just 1 line of code