Efficient Triton Kernels for LLM Training
Research project. A Memory solution for users, teams, and applications
Integrate cutting-edge LLM technology quickly and easily into your app
Secure, kernel-enforced sandbox CLI and SDKs for AI agents
FlashInfer: Kernel Library for LLM Serving
TT-NN operator library, and TT-Metalium low level kernel programming
A RWKV management and startup tool, full automation, only 8MB
Burn is a new comprehensive dynamic Deep Learning Framework
Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Training neural networks on Apple Neural Engine via APIs
Open source solution that can meet the requirements of workloads
FlashMLA: Efficient Multi-head Latent Attention Kernels
C++ library for high performance inference on NVIDIA GPUs
An experimental version of DeepSeek model
A Powerful Native Multimodal Model for Image Generation
Automate native Android apps with AI using accessibility APIs
AI memory OS for LLM and Agent systems
The Compute Library is a set of computer vision and machine learning
Low-latency AI inference engine optimized for mobile devices
Geometric deep learning extension library for PyTorch
Tool that provides interactive visualizations for large embeddings
oneAPI Deep Neural Network Library (oneDNN)
Deepnote is a drop-in replacement for Jupyter
How to optimize some algorithm in cuda
Deep and Machine Learning for Microscopy