nemotron-lightning-3.5-30b-a3b
NVIDIA · 1 configuration
Tariff: Requesty · Source: models-dev · USD per million tokens
What does more thinking buy you?
Compare this model’s measured thinking configurations.
No pair of measurements available
Change the filters or pick another benchmark. Missing values are not estimated.
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
This model is in the catalog but has no imported benchmarks. Capabilities are not estimated.