nemotron-3-ultra-nvfp4
NVIDIA · 1 configuration
Tariff: Requesty · Source: models-dev · USD per million tokens
What does more thinking buy you?
Compare this model’s measured thinking configurations.
No pair of measurements available
Change the filters or pick another benchmark. Missing values are not estimated.
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.
This model is in the catalog but has no imported benchmarks. Capabilities are not estimated.