DeepSeek 4 Flash local inference engine for Metal
Port of OpenAI's Whisper model in C/C++
Run Local LLMs on Any Device. Open-source
Containerized automation engine for programmable CI/CD workflows
C++-based high-performance parallel environment execution engine
The easiest way to use Ollama in .NET
Fast LLM speculative inference server for consumer hardware
Low-latency machine code generation
Our first fully AI generated deep learning system
Declarative way to run AI models in React Native on device
ByteHook is an Android PLT hook library
TT-NN operator library, and TT-Metalium low level kernel programming
Terminal-native coding agent powered by local LLMs
Training neural networks on Apple Neural Engine via APIs
Fastest LLM inference runtime for Apple Silicon
Pre-trained Deep Learning models and demos
FlashMLA: Efficient Multi-head Latent Attention Kernels
fast C++ library for GPU linear algebra & scientific computing
C++ Statistical ToolKit
Rust language bindings for TensorFlow
.NET SDK for processing phone calls and SMS through the VoiceShot API.