NVIDIA Releases Nemotron 3.5 Lightning: 30B MoE with Only 3B Active Parameters

NVIDIA AI has released Nemotron 3.5 Lightning, a Mixture-of-Experts model with 30 billion total parameters but only 3 billion active at inference time, paired with the NeMo Switchyard model router for intelligent request routing. The low active-parameter count means inference costs are dramatically reduced compared to dense models of equivalent capacity, making it viable for production deployments where latency and cost matter. NVIDIA also released the NeMo Switchyard router alongside it, which lets developers automatically route requests to the most appropriate model in a fleet — a key primitive for multi-model agentic systems. For developers building with NVIDIA's ecosystem, this is a direct path to running capable reasoning at dense-model quality with MoE-level efficiency. The combination of a strong open MoE and a production-ready router makes this a meaningful infrastructure upgrade for teams running self-hosted inference.
Read original source ↗Part of the 2026-08-13 digest→