NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 is an open large language model optimized for efficient autonomous agents, sub-agent deployments, and local inference. It uses a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selected attention layers, with 30B total parameters but only 3B active during inference. The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning. Its NVFP4 quantization reduces deployment requirements while targeting NVIDIA hardware ranging from DGX Spark and RTX 5090 systems to H100, H200, and GB200 accelerators. NVIDIA also provides DSpark, Multi-Token Prediction, and DFlash speculative decoding methods to accelerate text generation. It supports English, coding languages, Spanish, French, German, Italian, and Japanese, and is intended for commercially deployable AI applications. The model can run on a single DGX Spark or H100 and integrates with inference frameworks including vLLM.

Features

  • 30B total parameters with only 3B active
  • Hybrid Mamba-2, MoE, and attention architecture
  • Up to 1M-token context length
  • NVFP4 quantization for memory-efficient deployment
  • DSpark, MTP, and DFlash speculative decoding support
  • Single-GPU deployment on DGX Spark or H100
  • Multilingual support plus programming languages
  • Optimized for autonomous agents, sub-agents, and local AI

Project Samples

Project Activity

See All Activity >

Categories

AI Models

Follow Nemotron 3.5 Lightning

Nemotron 3.5 Lightning Web Site

Other Useful Business Software
Fully Managed MySQL, PostgreSQL, and SQL Server Icon
Fully Managed MySQL, PostgreSQL, and SQL Server

Automatic backups, patching, replication, and failover. Focus on your app, not your database.

Cloud SQL handles your database ops end to end, so you can focus on your app.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Nemotron 3.5 Lightning!

Additional Project Details

Registered

2026-08-13