Core capabilities
- Text, image, video, audio and PDF understanding
- 1,048,576 input tokens and 65,536 output tokens
- Thinking, code execution, search grounding and URL context
- Function calling, structured output, file search and computer use
Google / stable
A speed-intelligence balance for agentic, coding and complex multimodal spatial tasks.
Its large context and native multimodality suit production pipelines mixing documents, video and tools. Google positions 3.6 Flash for complex agentic, coding and multimodal work with fewer turns and tool calls. It supports a 1,048,576-token input limit, 65,536-token output and production GA status.
Use the stable model ID rather than a latest alias. Migrate to thinking_level, remove deprecated sampling parameters, validate conversation-turn rules, and log tool loops so reduced-turn behavior can be measured in production.
Stable, preview and latest aliases behave differently; pin a stable model for production. Gemini 3.x changes generation controls: deprecated sampling fields and old thinking-budget patterns should be removed during migration. Explicit design direction still matters for visual styling tasks.
Specifications and positioning checked against provider documentation.
Open official source↗