Z

Zhipu: glm-5.3-flash-thinking:free

glm-5.3-flash-thinking:free
1M context943.7K outReasoningToolsParallel toolsVisionVideoCacheStructured
Released Aug 26, 2026Updated Aug 26, 2026

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Mode chatTokenizer OtherQuantization fp8zai-org/GLM-5.3-Flash

All providers for this model are busy right now

Every upstream provider has hit its rate limit. The model comes back automatically once limits lift, usually within hours. Try again in a little while or switch to another model.

Request this model on Discord

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Compatible endpoints openaiVendor Zhipu

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providers-
include_reasoningAll providers-
logit_biasSome providers-
logprobsSome providers-
max_tokensAll providers-
min_pSome providers-
parallel_tool_callsSome providers-
presence_penaltyAll providers-
reasoningAll providers-
reasoning_effortAll providers-
repetition_penaltyAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers1
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_logprobsSome providers-
top_pAll providers0.95

Frequently asked questions

Similar models