D

DeepSeek: deepseek-v4.1-flash

deepseek-v4.1-flash
1.0M context384K outReasoningToolsVisionCacheStructured
Released Sep 10, 2026Updated Sep 10, 2026

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family. DeepSeek reports that it exceeds V4 Pro on performance, speed, and task completion time, so it sits ahead of the previous flagship rather than beneath it. It is suited for coding, reasoning, and agentic workflows, and is particularly strong at long-horizon tasks that must run to completion across many steps.

Tokenizer DeepSeekQuantization fp8deepseek-ai/DeepSeek-V4.1-Flash

Pricing

Input price
$0.07/ 1M tokens
Output price
$0.21/ 1M tokens
Compatible endpoints openaiVendor DeepSeek

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltyAll providers-
include_reasoningAll providers-
logit_biasSome providers-
logprobsAll providers-
max_tokensAll providers-
min_pSome providers-
presence_penaltyAll providers-
reasoningAll providers-
reasoning_effortAll providers-
repetition_penaltyAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providers-
tool_choiceAll providers-
toolsAll providers-
top_kAll providers-
top_logprobsAll providers-
top_pAll providers-

Frequently asked questions

Similar models