N

NVIDIA: nemotron-3-ultra-550b-a55b:free

nemotron-3-ultra-550b-a55b:free
262.1K context32.8K outReasoningToolsCacheStructured
Released Jun 4, 2026Updated Jun 18, 2026

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it supports text input and output with a context window of up to 1M tokens. It is suited for long-running agentic workflows, including agent orchestration, coding agents, deep research, and complex enterprise tasks. It is particularly strong at multi-step reasoning and planning, with high-throughput inference designed for high-volume agent pipelines. It is part of the NVIDIA Nemotron family of open models for agentic AI.

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 512.3K tokensCompatible endpoints openaiVendor NVIDIA

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providersNot sent by default
include_reasoningAll providers-
logit_biasAll providers-
max_tokensAll providers-
min_pAll providers-
presence_penaltyAll providersNot sent by default
reasoningAll providers-
reasoning_effortAll providers-
repetition_penaltyAll providersNot sent by default
response_formatAll providers-
seedSome providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers1
tool_choiceAll providers-
toolsAll providers-
top_kAll providersNot sent by default
top_pAll providers0.95

Frequently asked questions

Similar models