G
Google: gemini-2.5-flash-lite
gemini-2.5-flash-lite
1M context65.5K outReasoningToolsParallel toolsVisionAudio inVideoFilesCacheStructuredWeb searchURL contextSystem msg
Released Jul 22, 2025Knowledge cutoff Jan 2025Updated Sep 22, 2026
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.
Mode chatTokenizer GeminiDeprecation Oct 2026Expiration Oct 2026
Pricing
Input price
$0.09/ 1M tokens7% off
Output price
$0.37/ 1M tokens7% off
Compatible endpoints openaiVendor Google
Uptime
Performance
Loading performance data...
Usage & Ranking
Loading usage...
Supported parameters
All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.
| Parameter | Providers | Default |
|---|---|---|
| frequency_penalty | Some providers | Not sent by default |
| include_reasoning | All providers | - |
| max_tokens | All providers | - |
| reasoning | All providers | - |
| response_format | All providers | - |
| seed | All providers | - |
| stop | All providers | - |
| structured_outputs | All providers | - |
| temperature | All providers | Not sent by default |
| tool_choice | All providers | - |
| tools | All providers | - |
| top_p | All providers | Not sent by default |