NVIDIA: nemotron-3-ultra-550b-a55b-free:free
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it supports text input and output with a context window of up to 1M tokens. It is suited for long-running agentic workflows, including agent orchestration, coding agents, deep research, and complex enterprise tasks. It is particularly strong at multi-step reasoning and planning, with high-throughput inference designed for high-volume agent pipelines. It is part of the NVIDIA Nemotron family of open models for agentic AI.
All providers for this model are busy right now
Every upstream provider has hit its rate limit. The model comes back automatically once limits lift, usually within hours. Try again in a little while or switch to another model.
Request this model on DiscordPricing
Uptime
Performance
Usage & Ranking
Supported parameters
All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.
| Parameter | Providers | Default |
|---|---|---|
| frequency_penalty | All providers | Not sent by default |
| include_reasoning | All providers | - |
| logit_bias | Some providers | - |
| max_tokens | All providers | - |
| min_p | Some providers | - |
| presence_penalty | All providers | Not sent by default |
| reasoning | All providers | - |
| reasoning_effort | All providers | - |
| repetition_penalty | Some providers | Not sent by default |
| response_format | All providers | - |
| seed | Some providers | - |
| stop | All providers | - |
| structured_outputs | Some providers | - |
| temperature | All providers | 1 |
| tool_choice | All providers | - |
| tools | All providers | - |
| top_k | All providers | Not sent by default |
| top_p | All providers | 0.95 |