Ling 3.0 Tiny
Lightweight MoE model for local reasoning, coding, and AI agents
...The model supports both fast responses and configurable multi-step thinking, covering general agents, coding, mathematics, scientific reasoning, and instruction following. It is specifically optimized for local and resource-constrained deployment and has been validated on NVIDIA DGX Spark, Apple Silicon MacBooks, and Mac mini systems. FP8 testing reached around 100–105 tokens/s on DGX Spark and 86–90 tokens/s on an M4 Pro MacBook. BF16, FP8, and INT4 weights are available, while deployment options include SGLang, vLLM, and experimental Ollama support on Apple Silicon.