Alibaba: qwen3.6-35b-a3b:free
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers, enabling efficient inference at a fraction of the compute cost. The model supports a 262K token native context window (extensible to 1M via YaRN) and accepts text, image, and video inputs. It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations, function calling, and structured output. Released under the Apache 2.0 license.
Pricing
Uptime
Performance
Usage & Ranking
Supported parameters
All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.
| Parameter | Providers | Default |
|---|---|---|
| frequency_penalty | All providers | - |
| include_reasoning | All providers | - |
| logit_bias | Some providers | - |
| logprobs | All providers | - |
| max_tokens | All providers | - |
| min_p | Some providers | - |
| presence_penalty | All providers | - |
| reasoning | All providers | - |
| repetition_penalty | All providers | - |
| response_format | All providers | - |
| seed | All providers | - |
| stop | All providers | - |
| structured_outputs | All providers | - |
| temperature | All providers | 1 |
| tool_choice | All providers | - |
| tools | All providers | - |
| top_k | All providers | 20 |
| top_logprobs | All providers | - |
| top_p | All providers | 0.95 |