Android GPU Inspector
How to optimize some algorithm in cuda
A simple Minecraft modpack focusing on performance and graphics
Performance meets Productivity
Development repository for the Triton language and compiler
A Python framework for accelerated simulation, data generation
OpenLIT is an open-source LLM Observability tool
Pruna is a model optimization framework built for developers
Performance-optimized AI inference on your GPUs
State-of-the-art Parameter-Efficient Fine-Tuning
High-performance Kimi Delta Attention kernels
Fast and memory-efficient exact attention
AI agents running research on single-GPU nanochat training
The CUDA target for Numba
Meridian is an MMM framework
GPU accelerated decision optimization
An open-source, GPU-accelerated physics simulation engine
Ongoing research training transformer models at scale
Supercharge Your LLM with the Fastest KV Cache Layer
Python inference and LoRA trainer package for the LTX-2 audio–video
Bridging Reasoning and Action Prediction
Faster Whisper transcription with CTranslate2
The Modular Platform (includes MAX & Mojo)
Unified KV Cache Compression Methods for Auto-Regressive Models
Low-latency REST API for serving text-embeddings