Z

Zhipu: glm-4.5-flash:free

glm-4.5-flash:free
131.1K context32K outReasoningToolsParallel toolsCacheStructured
Released Jul 25, 2025Knowledge cutoff Dec 2024Updated Sep 14, 2026

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid inference mode with two options, a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for instant responses. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Mode chatTokenizer OtherQuantization fp8Expiration Dec 2026zai-org/GLM-4.5

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 131.1K tokensCompatible endpoints openaiVendor Zhipu

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltySome providersNot sent by default
include_reasoningAll providers-
max_tokensAll providers-
reasoningAll providers-
response_formatAll providers-
temperatureAll providers0.75
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_pAll providersNot sent by default

Frequently asked questions

Similar models