Gemini 3.5 Flash
Google's most capable Flash model, tuned for the agentic era: sub-agent deployment, multi-step workflows, and rapid coding iterations at scale. Supports search grounding, function calling, structured outputs, and computer use in preview. Note that it is pricier than the Gemini 3 Flash preview it replaced.
Pricing checked against provider documentation on . How we verify
gemini-3.5-flash reasoning Input
$1.50
/ 1M tokens
Output
$9.00
/ 1M tokens
Context
1.05M
1,048,576 tokens
Max output
66K
tokens per request
Pricing caveat: Flat rate at any context length, unlike Gemini 3.1 Pro's 200K cliff.
What it costs in practice
A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $33.00/month on Gemini 3.5 Flash. Routing that same volume through the batch API brings it to $16.50. The next cheaper option, GLM-5.2 , would cost $22.80.
Where Gemini 3.5 Flash fits
- Sub-agent fleets and orchestration
- Fast coding iteration loops
- Search-grounded answers
- Multimodal ingestion at scale
Specifications
- API model ID
- gemini-3.5-flash
- Provider
- Model family
- Gemini 3.5
- Released
- May 2026
- Knowledge cutoff
- January 2025
- Context window
- 1,048,576 tokens
- Max output
- 65,536 tokens
- Native reasoning
- Yes
- Relative latency
- low
- Input modalities
- text, image, audio, video, pdf
- Output modalities
- text
- Open weights
- No
- Cached input
- —
- Batch discount
- 50%
Frequently asked
How much does Gemini 3.5 Flash cost?
Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens. Batch processing is 50% cheaper.
What is Gemini 3.5 Flash's context window?
Gemini 3.5 Flash accepts up to 1,048,576 tokens of context and can generate up to 65,536 output tokens per request.
Is Gemini 3.5 Flash a reasoning model?
Yes. Gemini 3.5 Flash performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.
What is Gemini 3.5 Flash best for?
Agentic loops and sub-agent fleets. Google's most capable Flash model, tuned for the agentic era: sub-agent deployment, multi-step workflows, and rapid coding iterations at scale. Supports search grounding, function calling, structured outputs, and computer use in preview. Note that it is pricier than the Gemini 3 Flash preview it replaced.