Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Performance meets Productivity
The CUDA target for Numba
How to optimize some algorithm in cuda
A NumPy-compatible array library accelerated by CUDA
Machine Learning Containers for NVIDIA Jetson and JetPack-L4T
The best AI Aimbot for Fortnite, Valorant, CS2, R6, COD, Apex, & more
High-performance Kimi Delta Attention kernels
Solve puzzles. Learn CUDA
A Python framework for accelerated simulation, data generation
A high-throughput and memory-efficient inference and serving engine
Development repository for the Triton language and compiler
Our first fully AI generated deep learning system
Rembg is a tool to remove images background
Prevent PyTorch's `CUDA error: out of memory` in just 1 line of code
Self-host the powerful Chatterbox TTS model
Stable Diffusion built-in to Blender
High-Resolution Image Synthesis with Latent Diffusion Models
A lightweight vLLM implementation built from scratch
Apple Silicon (MLX) port of Karpathy's autoresearch
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
Geometric deep learning extension library for PyTorch
Package and deploy machine learning models using Docker containers
Jittor is a high-performance deep learning framework