NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 is an open large language model optimized for efficient autonomous agents, sub-agent deployments, and local inference. It uses a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selected attention layers, with 30B total parameters but only 3B active during inference. The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning. Its NVFP4 quantization reduces deployment requirements while targeting NVIDIA hardware ranging from DGX Spark and RTX 5090 systems to H100, H200, and GB200 accelerators. NVIDIA also provides DSpark, Multi-Token Prediction, and DFlash speculative decoding methods to accelerate text generation. It supports English, coding languages, Spanish, French, German, Italian, and Japanese, and is intended for commercially deployable AI applications. The model can run on a single DGX Spark or H100 and integrates with inference frameworks including vLLM.
Features
- 30B total parameters with only 3B active
- Hybrid Mamba-2, MoE, and attention architecture
- Up to 1M-token context length
- NVFP4 quantization for memory-efficient deployment
- DSpark, MTP, and DFlash speculative decoding support
- Single-GPU deployment on DGX Spark or H100
- Multilingual support plus programming languages
- Optimized for autonomous agents, sub-agents, and local AI