N

NVIDIA: nemotron-3-super-120b-a12b-free:free

nemotron-3-super-120b-a12b-free:free
262.1K context16.4K outReasoningToolsStructured
Released Mar 11, 2026Updated Sep 4, 2026

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models. The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud.

Mode chatTokenizer OtherQuantization bf16nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

All providers for this model are busy right now

Every upstream provider has hit its rate limit. The model comes back automatically once limits lift, usually within hours. Try again in a little while or switch to another model.

Request this model on Discord

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 262.1K tokensCompatible endpoints -Vendor NVIDIA

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providersNot sent by default
include_reasoningAll providers-
logit_biasAll providers-
max_tokensAll providers-
min_pAll providers-
presence_penaltyAll providersNot sent by default
reasoningAll providers-
reasoning_effortAll providers-
repetition_penaltyAll providersNot sent by default
response_formatAll providers-
seedAll providers-
stopAll providers-
temperatureAll providers1
tool_choiceAll providers-
toolsAll providers-
top_kAll providersNot sent by default
top_pAll providers0.95

Frequently asked questions

Similar models