Laya-MLX is an independent MLX implementation of Laya’s typed decision models for Apple Silicon Macs. It performs structured choice, score, and probability decisions without token-by-token text generation. Inference runs fully locally after model weights are downloaded and does not require PyTorch, Transformers, or a cloud API. The runtime supports English, multilingual, and typed-decision Laya checkpoints while preserving their original calibration and output formats. Its implementation moves the encoder, decision transformer, scoring head, and action head into MLX. The project reports short-decision latency in the single-digit to low-teens millisecond range on an M3 Max, depending on the checkpoint. It also includes routing utilities, demos, validation tests, benchmarks, and a Python API.
Features
- Native Apple Silicon MLX inference
- Choice, score, and probability decisions
- Fully local execution
- English and multilingual checkpoints
- Python API and model routing
- Benchmarks, validation tests, and demos