Find a model

Search models and open their dedicated profile. Use the arrow keys to navigate and Enter to open.

nemotron-lightning-3.5-30b-a3b

NVIDIA · 1 configuration

Intelligence—
Input / 1M$0.05
Output / 1M$0.20
Context262.1K

Tariff: Requesty · Source: models-dev · USD per million tokens

Thinking configuration

What does more thinking buy you?

Compare this model’s measured thinking configurations.

◌

No pair of measurements available

Change the filters or pick another benchmark. Missing values are not estimated.

Selected: thinking default · Click a point to switch configuration

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.

This model is in the catalog but has no imported benchmarks. Capabilities are not estimated.

Specifications from the catalog

Inputtext
Outputtext
Tool callingYes
Structured outputYes
SpeedNot available
First tokenNot available
Maximum output262.1K
Catalog listing date2026-08-15