Inkling-SmallThinking Machines Lab
|
Olmo 3Ai2
|
|||||
Related Products
|
||||||
About
Inkling-Small is an efficient model that offers performance comparable to Inkling at a quarter of its size. It is a Mixture-of-Experts transformer with 276 billion total parameters and 12 billion active parameters, trained on NVIDIA GB300 NVL72 systems. It supports native reasoning across text, images, and audio, variable thinking effort, and context windows of up to one million tokens. Users adjust reasoning effort from minimal to extra high to balance performance and compute. Improved pre-training data, post-training with on-policy distillation from Inkling, and extended agentic coding reinforcement learning helped Inkling-Small surpass its larger counterpart on reasoning and coding benchmarks. It performs well in coding and tool-use harnesses, exceeds 80% on SWE-bench Verified, and combines strong reasoning with efficient output. Its encoder-free multimodal architecture processes audio as dMel spectrograms and images as 40-by-40-pixel patches alongside text tokens.
|
About
Olmo 3 is a fully open model family spanning 7 billion and 32 billion parameter variants that delivers not only high-performing base, reasoning, instruction, and reinforcement-learning models, but also exposure of the entire model flow, including raw training data, intermediate checkpoints, training code, long-context support (65,536 token window), and provenance tooling. Starting with the Dolma 3 dataset (≈9 trillion tokens) and its disciplined mix of web text, scientific PDFs, code, and long-form documents, the pre-training, mid-training, and long-context phases shape the base models, which are then post-trained via supervised fine-tuning, direct preference optimisation, and RL with verifiable rewards to yield the Think and Instruct variants. The 32 B Think model is described as the strongest fully open reasoning model to date, competitively close to closed-weight peers in math, code, and complex reasoning.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Developers, AI agent builders, software engineering teams, research teams, enterprise AI teams, multimodal application developers, coding assistant builders, tool-use workflow teams, and organizations that need efficient reasoning, long-context processing, text-image-audio understanding, adjustable thinking effort, coding performance, and scalable Mixture-of-Experts inference
|
Audience
AI researchers, developers and enterprises needing a tool offering foundation models to inspect, fine-tune or deploy with full provenance and auditability
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
$0.30 per million input tokens
$0.30 per million input tokens and $1.20 per million output tokens
Free Version
Free Trial
|
Pricing
Free
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Pros & Cons from Real UsersPros
Cons
|
||||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationThinking Machines Lab
Founded: 2025
United States
thinkingmachines.ai/news/inkling-small/
|
Company InformationAi2
Founded: 2014
United States
allenai.org/blog/olmo3
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
Model Context Protocol (MCP)
Tinker
|
||||||
|
|
|