G

Google: gemini-2.5-flash-lite

gemini-2.5-flash-lite
1M context65.5K outReasoningToolsParallel toolsVisionAudio inVideoFilesCacheStructuredWeb searchURL contextSystem msg
Released Jul 22, 2025Knowledge cutoff Jan 2025Updated Sep 22, 2026

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.

Mode chatTokenizer GeminiDeprecation Oct 2026Expiration Oct 2026

Pricing

Input price
$0.09/ 1M tokens7% off
Output price
$0.37/ 1M tokens7% off
Compatible endpoints openaiVendor Google

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltySome providersNot sent by default
include_reasoningAll providers-
max_tokensAll providers-
reasoningAll providers-
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providersNot sent by default
tool_choiceAll providers-
toolsAll providers-
top_pAll providersNot sent by default

Frequently asked questions

Similar models