A

Alibaba: qwen3-32b:free

qwen3-32b:free
131.1K context128K outReasoningToolsStructured
Released Apr 28, 2025Knowledge cutoff Mar 2025Updated Jun 18, 2026

A flagship dense model in the Qwen3 series, released by Alibaba Cloud on April 29, 2025. Featuring 32.8 billion parameters and a 64-layer Transformer architecture, it stands as a high-performance "mid-size" pillar of the Qwen family. It natively supports Dual-Mode (Thinking/Non-Thinking) switching, delivering logical reasoning depth comparable to previous-generation 72B or even 110B models within a 30B-scale footprint. With exceptional scores in AIME 2025 and LiveCodeBench, it is a premier choice for developers seeking a balance between inference speed and cognitive rigor.

Mode chatTokenizer Qwen3Quantization fp8Qwen/Qwen3-32B

All providers for this model are busy right now

Every upstream provider has hit its rate limit. The model comes back automatically once limits lift, usually within hours. Try again in a little while or switch to another model.

Request this model on Discord

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 131.1K tokensCompatible endpoints openaiVendor Alibaba

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providers-
include_reasoningAll providers-
logit_biasAll providers-
max_tokensAll providers-
min_pAll providers-
presence_penaltyAll providers-
reasoningAll providers-
repetition_penaltyAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers-
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_pAll providers-

Frequently asked questions

Similar models