A

Alibaba: qwen3.6-35b-a3b:free

qwen3.6-35b-a3b:free
262.1K context16.4K outReasoningToolsVisionVideoCacheStructured
Released Apr 27, 2026Updated Jun 18, 2026

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers, enabling efficient inference at a fraction of the compute cost. The model supports a 262K token native context window (extensible to 1M via YaRN) and accepts text, image, and video inputs. It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations, function calling, and structured output. Released under the Apache 2.0 license.

Mode chatTokenizer QwenQuantization fp8Qwen/Qwen3.6-35B-A3B

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 262.1K tokensCompatible endpoints openaiVendor Alibaba

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providers-
include_reasoningAll providers-
logit_biasSome providers-
logprobsAll providers-
max_tokensAll providers-
min_pSome providers-
presence_penaltyAll providers-
reasoningAll providers-
repetition_penaltyAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers1
tool_choiceAll providers-
toolsAll providers-
top_kAll providers20
top_logprobsAll providers-
top_pAll providers0.95

Frequently asked questions

Similar models