Core capabilities
- High-throughput multimodal document parsing
- Code execution, function calling and structured outputs
- Caching, file search, search grounding and URL context
- Batch, flex and priority inference options
Google / stable
The fastest, lowest-cost Gemini 3.5 variant, optimized for high-throughput execution.
Well suited to extraction, structured JSON, data analysis and autonomous subagents with clear scopes. Flash-Lite is Google's fastest, lowest-cost 3.5 model for document parsing, structured extraction and subagent execution. It accepts text, image, video, audio and PDF, with 1,048,576 input and 65,536 output limits.
Use minimal thinking for extraction and simple automation; test higher thinking for planning-heavy subagents. Enforce JSON schemas, batch predictable workloads and monitor document-level accuracy rather than token-level fluency.
Do not force complex creative or high-stakes judgment onto the light tier purely for price. Computer Use support differs by API surface and model documentation, so confirm the exact endpoint rather than assuming feature parity with 3.6 Flash. Raise thinking level only when added planning improves measured quality.
Specifications and positioning checked against provider documentation.
Open official source↗