Z
Zhipu: glm-5.3-flash-thinking:free
glm-5.3-flash-thinking:free
1M context943.7K outReasoningToolsParallel toolsVisionVideoCacheStructured
Released Aug 26, 2026Updated Aug 26, 2026
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
All providers for this model are busy right now
Every upstream provider has hit its rate limit. The model comes back automatically once limits lift, usually within hours. Try again in a little while or switch to another model.
Request this model on DiscordPricing
Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Compatible endpoints openaiVendor Zhipu
Uptime
Performance
Loading performance data...
Usage & Ranking
Loading usage...
Supported parameters
All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.
| Parameter | Providers | Default |
|---|---|---|
| frequency_penalty | All providers | - |
| include_reasoning | All providers | - |
| logit_bias | Some providers | - |
| logprobs | Some providers | - |
| max_tokens | All providers | - |
| min_p | Some providers | - |
| parallel_tool_calls | Some providers | - |
| presence_penalty | All providers | - |
| reasoning | All providers | - |
| reasoning_effort | All providers | - |
| repetition_penalty | All providers | - |
| response_format | All providers | - |
| seed | All providers | - |
| stop | All providers | - |
| structured_outputs | All providers | - |
| temperature | All providers | 1 |
| tool_choice | All providers | - |
| tools | All providers | - |
| top_k | All providers | - |
| top_logprobs | Some providers | - |
| top_p | All providers | 0.95 |