A

Alibaba: qwen3.5-flash:free

qwen3.5-flash:free
1M context65.5K outReasoningToolsVisionVideoStructured
Released Feb 25, 2026Updated Aug 6, 2026

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.

Mode chatTokenizer Qwen3

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Compatible endpoints openaiVendor Alibaba

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providersNot sent by default
include_reasoningAll providers-
max_tokensAll providers-
presence_penaltyAll providers-
reasoningAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providersNot sent by default
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_pAll providersNot sent by default

Frequently asked questions

Similar models