G
Google: gemini-3.5-flash:free
gemini-3.5-flash:free
1M context65.5K outReasoningToolsParallel toolsVisionAudio inVideoFilesCacheStructuredWeb searchURL contextStreamingSystem msg
Released May 19, 2026Knowledge cutoff Jan 2025Updated Jun 20, 2026
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution loops, supporting text, image, video, audio, and PDF inputs. Defaults to medium thinking effort for faster and more cost-efficient responses, with full support for thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs.
Mode chatTokenizer GeminiDeprecation May 2027
Pricing
Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Compatible endpoints openaiVendor Google
Uptime
Performance
Loading performance data...
Usage & Ranking
Loading usage...
Supported parameters
All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.
| Parameter | Providers | Default |
|---|---|---|
| frequency_penalty | Some providers | Not sent by default |
| include_reasoning | All providers | - |
| max_tokens | All providers | - |
| presence_penalty | Some providers | Not sent by default |
| reasoning | All providers | - |
| reasoning_effort | All providers | - |
| repetition_penalty | Some providers | Not sent by default |
| response_format | All providers | - |
| seed | All providers | - |
| stop | All providers | - |
| structured_outputs | All providers | - |
| temperature | All providers | Not sent by default |
| tool_choice | All providers | - |
| tools | All providers | - |
| top_k | Some providers | Not sent by default |
| top_p | All providers | Not sent by default |