Z

Zhipu: glm-4.7-flash-heretic:free

glm-4.7-flash-heretic:free
200K context131.1K outReasoningToolsCacheStructured
Released Jan 19, 2026Updated Jun 20, 2026

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Mode chatTokenizer Otherzai-org/GLM-4.7-Flash

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Context window 204.8K tokensCompatible endpoints openaiVendor Zhipu

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providersNot sent by default
include_reasoningAll providers-
logit_biasAll providers-
logprobsSome providers-
max_tokensAll providers-
min_pAll providers-
presence_penaltyAll providers-
reasoningAll providers-
repetition_penaltyAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers1
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_logprobsSome providers-
top_pAll providers0.95

Frequently asked questions

Similar models