Nemotron 3.5 Lightning

Tiny open 30B-A3B model with the lowest first-token latency in the index.

Nemotron 3.5 Lightning is a 30B-parameter, 3B-active open model tuned for very low latency and high throughput.

Provider
NVIDIA
Type
Language model
Released
Aug 11, 2026
Context window
262K tokens
Max output
236K tokens
Price
$0.070 input / $0.20 output per 1M tokens
Input
text
Output
text
Open weights
Yes
License
NVIDIA Open Model License

Best for

Strengths

More from NVIDIA

Sources