Nvidia's Nemotron 3.5 Lightning: 30B Model Runs on a Single GPU
Nvidia Nemotron 3.5 Lightning
Nvidia unveiled Nemotron 3.5 Lightning, a 30B-parameter MoE hybrid model with 3B active parameters, optimized for single-GPU deployment on DGX Spark or H100. It supports up to 1M token context and features speculative decoding via DSpark, DFlash, or MTP. The model excels in long-running agent tasks, with strong SWE-bench Verified and GPQA Diamond scores, and is available under the OpenMDW-1.1 license.
The model has 3B active parameters and 30B parameters in total.