Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
How to optimize some algorithm in cuda
Machine Learning Containers for NVIDIA Jetson and JetPack-L4T
Solve puzzles. Learn CUDA
A high-throughput and memory-efficient inference and serving engine
Our first fully AI generated deep learning system
Prevent PyTorch's `CUDA error: out of memory` in just 1 line of code
Self-host the powerful Chatterbox TTS model
Stable Diffusion built-in to Blender
High-Resolution Image Synthesis with Latent Diffusion Models
A lightweight vLLM implementation built from scratch
Apple Silicon (MLX) port of Karpathy's autoresearch
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
Geometric deep learning extension library for PyTorch
Package and deploy machine learning models using Docker containers
Jittor is a high-performance deep learning framework
Fast and memory-efficient exact attention
Interface for OuteTTS models
Universal LLM Deployment Engine with ML Compilation
A Python library for learning and evaluating knowledge graph embedding
Fast Python collaborative filtering for implicit feedback datasets
An experimental version of DeepSeek model
Generate audiobooks from e-books
A simple native web interface that uses ChatTTS to synthesize text