Strata is an open-source local inference system for running Qwen3.8-Flash-Next on consumer Windows and Linux PCs. It is designed to make a very large mixture-of-experts model practical with one NVIDIA GPU, system RAM, and local storage instead of server-class hardware. Frequently used experts remain on the GPU while other model components are shared across RAM, CPU, and SSD resources. Speculative decoding uses a smaller helper component to propose tokens that the main model verifies in batches. Long prompts are processed in large blocks to improve ingestion speed. Strata provides a local chat interface and exposes OpenAI- and Anthropic-compatible APIs for other applications. Optional image input and several quantized model sizes allow users to balance memory use, speed, and quality.
Features
- Consumer-hardware large-model inference
- GPU, CPU, RAM, and SSD workload sharing
- Speculative decoding acceleration
- Fast long-context prompt processing
- OpenAI- and Anthropic-compatible local APIs
- Optional image input and quantized models