Research project. A Memory solution for users, teams, and applications
Efficient Triton Kernels for LLM Training
Integrate cutting-edge LLM technology quickly and easily into your app
FlashInfer: Kernel Library for LLM Serving
TT-NN operator library, and TT-Metalium low level kernel programming
Secure, kernel-enforced sandbox CLI and SDKs for AI agents
Burn is a new comprehensive dynamic Deep Learning Framework
Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Training neural networks on Apple Neural Engine via APIs
A RWKV management and startup tool, full automation, only 8MB
Deepnote is a drop-in replacement for Jupyter
Open source solution that can meet the requirements of workloads
The easiest way to use Ollama in .NET
An experimental version of DeepSeek model
FlashMLA: Efficient Multi-head Latent Attention Kernels
Library for efficiently connecting and optimizing teams of AI agents
oneAPI Deep Neural Network Library (oneDNN)
Automate native Android apps with AI using accessibility APIs
Tool that provides interactive visualizations for large embeddings
Geometric deep learning extension library for PyTorch
A Powerful Native Multimodal Model for Image Generation
The Compute Library is a set of computer vision and machine learning
Pytorch Distributed native training library for LLMs/VLMs
C++ library for high performance inference on NVIDIA GPUs
Clean and efficient FP8 GEMM kernels with fine-grained scaling