N

NVIDIA: nemotron-3-super-120b-a12b-free:free

nemotron-3-super-120b-a12b-free:free
262.1K contesto16.4K outputRagionamentoStrumentiStrutturato
Rilasciato Mar 11, 2026Aggiornato Sep 4, 2026

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models. The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud.

Modalità chatTokenizer OtherQuantizzazione bf16nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

Tutti i provider di questo modello sono occupati al momento

Ogni provider a monte ha raggiunto il suo limite di velocità. Il modello torna automaticamente quando i limiti si allentano, di solito entro poche ore. Riprova tra poco o passa a un altro modello.

Richiedi questo modello su Discord

Prezzi

Prezzo di input
$0.00/ 1 M token
Prezzo di output
$0.00/ 1 M token
Finestra di contesto 262.1K tokenEndpoint compatibili -Provider NVIDIA

Disponibilità

Performance

Caricamento dati di performance...

Utilizzo e classifica

Caricamento utilizzo...

Parametri supportati

Tutti i provider = supportato da ogni upstream che serve questo modello. Alcuni provider = dipende dall'upstream che gestisce la richiesta. Predefinito = il valore inviato quando non lo imposti.

ParametroProviderPredefinito
frequency_penaltyTutti i providerNon inviato per impostazione predefinita
include_reasoningTutti i provider-
logit_biasTutti i provider-
max_tokensTutti i provider-
min_pTutti i provider-
presence_penaltyTutti i providerNon inviato per impostazione predefinita
reasoningTutti i provider-
reasoning_effortTutti i provider-
repetition_penaltyTutti i providerNon inviato per impostazione predefinita
response_formatTutti i provider-
seedTutti i provider-
stopTutti i provider-
temperatureTutti i provider1
tool_choiceTutti i provider-
toolsTutti i provider-
top_kTutti i providerNon inviato per impostazione predefinita
top_pTutti i provider0.95

Domande frequenti

Modelli simili