G

Google: gemini-3.5-flash:free

gemini-3.5-flash:free
1M context65.5K outReasoningToolsParallel toolsVisionAudio inVideoFilesCacheStructuredWeb searchURL contextStreamingSystem msg
Released May 19, 2026Knowledge cutoff Jan 2025Updated Jun 20, 2026

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution loops, supporting text, image, video, audio, and PDF inputs. Defaults to medium thinking effort for faster and more cost-efficient responses, with full support for thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs.

Mode chatTokenizer GeminiDeprecation May 2027

Pricing

Input price
$0.00/ 1M tokens
Output price
$0.00/ 1M tokens
Compatible endpoints openaiVendor Google

Uptime

Performance

Loading performance data...

Usage & Ranking

Loading usage...

Supported parameters

All providers = every upstream serving this model supports it. Some providers = depends on which upstream handles the request. Default = the value sent when you leave the parameter unset.

ParameterProvidersDefault
frequency_penaltySome providersNot sent by default
include_reasoningAll providers-
max_tokensAll providers-
presence_penaltySome providersNot sent by default
reasoningAll providers-
reasoning_effortAll providers-
repetition_penaltySome providersNot sent by default
response_formatAll providers-
seedAll providers-
stopAll providers-
structured_outputsAll providers-
temperatureAll providersNot sent by default
tool_choiceAll providers-
toolsAll providers-
top_kSome providersNot sent by default
top_pAll providersNot sent by default

Frequently asked questions

Similar models