Skip to main content

Gemini 3.5 Flash-Lite

GoogleGemini

Description

Gemini 3.5 Flash-Lite is Google's low-latency, cost-effective multimodal reasoning model for high-throughput agentic workflows, document processing, data extraction, translation, and classification. It supports text, image, video, audio, and PDF inputs, a 1 million-token context window, and a 65,536-token text output.

Release Date
2026-07-21
Parameters
Context Length
1.0M
Modalities
audio, image, pdf, text, video

Capability Radar

32
general
48
coding
84
reasoning
56
scienceest.
50
agents
80
multimodal

Science uses a reasoning proxy when dedicated science benchmarks are unavailable.

Rankings

Domain#RankScoreSource
Code Ranking101
69.0
AA
General Ranking150
60.0
AA
Science124
62.0
AA

Benchmark Scores (LLM Stats)

Agents

GDPval-AA1140.00 / 3000SR
OSWorld-Verified74.0%SR
SWE-Bench Pro54.2%SR
Terminal-Bench 2.154.0%SR
MLE-Bench39.2%SR

General

MRCR v2 (8-needle)21.3%SR

Multimodal

CharXiv-R76.5%SR

AA Evaluation Indices

Coding Index
49.3
Intelligence Index
36.5
Gpqa
0.8
Lcr
0.6
Terminalbench V2 1
0.5
Scicode
0.4
Hle
0.2
Tau Banking
0.2

LLM Stats Category Scores

Legal
100
Finance
100
Agents
100
Reasoning
100
General
100
Multimodal
80
Vision
80
Code
50
Tool Calling
50
Long Context
20

Pricing

Input Price$0.3 / 1M tokens
Output Price$2.5 / 1M tokens
Blended Price (3:1)$0.85 / 1M tokens
Cache Read Price$0.03 / 1M tokens

Speed

Tokens/sec400.3
Time to First Token7.27s
Time to Answer7.27s

Provider Price Ranking

Provider Price Ranking

4 providers

Cheapest: GoogleMost Expensive: Venice AI
ProviderInputOutput
1GooglePRIMARY
$0.3
$2.5
2OpenRouter
$0.3
$2.5
3Vercel AI Gateway
$0.3
$2.5
4Venice AI
$0.375
$3.125

Compare pricing across different API providers for this model.

External Sources