I

InclusionAI: ling-3.0-flash-vl:free

ling-3.0-flash-vl:free
262.1K context32.8K outReasoningToolsVideoCacheStructured
Released Sep 10, 2026Updated Sep 11, 2026

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

Mode chatTokenizer OtherQuantization fp16inclusionAI/Ling-3.0-flash-VL

All providers for this model are busy right now

Every upstream provider has hit its rate limit. The model comes back automatically once limits lift, usually within hours. Try again in a little while or switch to another model.

Request this model on Discord

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 262.1K tokensCompatible endpoints -Vendor InclusionAI

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providers-
include_reasoningAll providers-
logit_biasAll providers-
max_tokensAll providers-
min_pAll providers-
presence_penaltyAll providers-
reasoningAll providers-
repetition_penaltyAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers-
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_pAll providers-

Frequently asked questions

Similar models