Python-free Rust inference server
Fastest LLM inference runtime for Apple Silicon
High-performance inference server for text embeddings models API layer
Switchyard lets LLM applications route traffic across models
Ghost in your shell. Ante is a self-contained agent harness
Toolkits to create a human-in-the-loop approval layer
Fast ML inference & training for ONNX models in Rust
Fast and efficient unstructured data extraction
Rust async runtime based on io-uring
Package and deploy machine learning models using Docker containers
Command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, etc.