D

DeepSeek: deepseek-v4-flash:free

deepseek-v4-flash:free
1M context384K outReasoningToolsCacheStructured
Released Apr 24, 2026Updated Jun 20, 2026

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

Mode chatTokenizer DeepSeekDeprecation Feb 2028deepseek-ai/DeepSeek-V4-Flash

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Compatible endpoints openaiVendor DeepSeek

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providers-
include_reasoningAll providers-
logit_biasSome providers-
logprobsSome providers-
max_completion_tokensSome providers-
max_tokensAll providers-
min_pSome providers-
presence_penaltyAll providers-
reasoningAll providers-
reasoning_effortAll providers-
repetition_penaltySome providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers-
tool_choiceAll providers-
toolsAll providers-
top_aSome providers-
top_kAll providers-
top_logprobsSome providers-
top_pAll providers-

Frequently asked questions

Similar models