...The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning. Its NVFP4 quantization reduces deployment requirements while targeting NVIDIA hardware ranging from DGX Spark and RTX 5090 systems to H100, H200, and GB200 accelerators. NVIDIA also provides DSpark, Multi-Token Prediction, and DFlash speculative decoding methods to accelerate text generation. It supports English, coding languages, Spanish, French, German, Italian, and Japanese, and is intended for commercially deployable AI applications. The model can run on a single DGX Spark or H100 and integrates with inference frameworks including vLLM.