Android GPU Inspector
How to optimize some algorithm in cuda
A simple Minecraft modpack focusing on performance and graphics
Ongoing research training transformer models at scale
Development repository for the Triton language and compiler
OpenLIT is an open-source LLM Observability tool
The CUDA target for Numba
An open-source, GPU-accelerated physics simulation engine
Fast and memory-efficient exact attention
A Python framework for accelerated simulation, data generation
AI agents running research on single-GPU nanochat training
GPU accelerated decision optimization
Performance meets Productivity
High-performance Kimi Delta Attention kernels
Unified KV Cache Compression Methods for Auto-Regressive Models
Faster Whisper transcription with CTranslate2
Performance-optimized AI inference on your GPUs
Supercharge Your LLM with the Fastest KV Cache Layer
The Modular Platform (includes MAX & Mojo)
Large Language Model Text Generation Inference
Python inference and LoRA trainer package for the LTX-2 audio–video
Pruna is a model optimization framework built for developers
Bridging Reasoning and Action Prediction
State-of-the-art Parameter-Efficient Fine-Tuning
Meridian is an MMM framework