Nemotron 3.5 Lightning
Tiny open 30B-A3B model with the lowest first-token latency in the index.
Nemotron 3.5 Lightning is a 30B-parameter, 3B-active open model tuned for very low latency and high throughput.
- Provider
- NVIDIA
- Type
- Language model
- Released
- Aug 11, 2026
- Context window
- 262K tokens
- Max output
- 236K tokens
- Price
- $0.070 input / $0.20 output per 1M tokens
- Input
- text
- Output
- text
- Open weights
- Yes
- License
- NVIDIA Open Model License
Best for
- Real-time / low latency
- High volume / low cost
- Local & on-device
- Extraction & classification
Strengths
- 0.57s first token
- 300+ tok/s
More from NVIDIA
- Nemotron 3 Super — NVIDIA. Open hybrid Mamba-Transformer MoE for multi-agent applications.
- Nemotron 3 Ultra — NVIDIA. NVIDIA’s largest open model: hybrid Mamba-Transformer MoE, 55B active.
- Nemotron 3 Embed 1B — NVIDIA. Small open embedding model for high-throughput enterprise retrieval.
- Cosmos 3 — NVIDIA. NVIDIA’s open omni-modal world foundation model for physical AI.
- Isaac GR00T N1.7 — NVIDIA. Open 3B cross-embodiment VLA for humanoids and manipulators.
- Nemotron 3.5 Content Safety — NVIDIA. Compact 4B multimodal guardrail model for LLM and VLM traffic.