BaseRT is a local large language model inference runtime optimized for Apple Silicon computers. It accelerates model execution through hand-written Metal kernels and requires an M1 or newer Mac running macOS 14 or later. A unified command-line interface can download models from Hugging Face, convert checkpoints, launch chats, benchmark performance, and inspect model packages. Its server implements OpenAI-compatible chat, completion, embedding, transcription, tool-calling, and multimodal endpoints. The custom .base format supports affine quantization from Q2 through Q8, optional AWQ calibration, and signed model bundles. Stable C interfaces connect the engine with Python, Node.js, Rust, and Swift applications. The repository contains the open CLI, format specifications, bindings, documentation, and benchmarks, while the prebuilt inference engine uses a separate license.

Features

  • Metal-accelerated Apple Silicon inference
  • Unified model management CLI
  • OpenAI-compatible local API server
  • Text, image, and audio model support
  • Q2 through Q8 model quantization
  • Python, Node.js, Rust, and Swift bindings

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow BaseRT

BaseRT Web Site

Other Useful Business Software
Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of BaseRT!

Additional Project Details

Programming Language

Rust

Related Categories

Rust Artificial Intelligence Software

Registered

2026-07-23