Skip to main content

Gemma 3 4B Instruct

GoogleGemmaOpen WeightGemma · Commercial OK

Description

Gemma 3 4B is a 4-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

Release Date
2025-03-12
Parameters
4.0B
Context Length
131K
Modalities
image, text

Capability Radar

14
general
6
coding
22
reasoning
17
science
17
agents
70
multimodal

Rankings

Domain#RankScoreSource
Code Ranking529
5.0
AA
General Ranking546
14.0
AA
Multimodal Ranking69
36.0
LS
Science537
14.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Biology

GPQANYU + Cohere + Anthropic (2023)30.8%SR

Code

HumanEvalOpenAI (2021)71.3%SR
LiveCodeBench12.6%SR

Factuality

FACTS Grounding70.1%SR
SimpleQA4.0%SR

Finance

MMLU-Pro43.6%SR

General

IFEvalGoogle Research (2023)90.2%SR
Natural2Code70.3%SR
MBPP0.63 / 100SR
Global-MMLU-Lite54.5%SR
MMMU (val)48.8%SR
BIG-Bench Extra Hard11.0%SR

Image To Text

DocVQADocVQA (2020)75.8%SR
VQAv2 (val)62.4%SR
TextVQA57.8%SR

Language

BIG-Bench Hard72.2%SR
WMT24++46.8%SR
ECLeKTic4.6%SR

Math

GSM8k89.2%SR
MATH75.6%SR
MathVista-Mini50.0%SR
HiddenMath43.0%SR

Multimodal

AI2D74.8%SR
ChartQAMasry et al. (2022)68.8%SR
InfoVQA50.0%SR

Reasoning

Bird-SQL (dev)36.3%SR

AA Evaluation Indices

(Artificial Analysis)
Math Index(Artificial Analysis)
12.7
Coding Index(Artificial Analysis)
2.7
Intelligence Index(Artificial Analysis)
1.0
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
0.8
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
0.4
Gpqa(NYU + Cohere + Anthropic (2023))
0.3
Ifbench(Google Research (2023))
0.3
Aime 25(MAA (Mathematical Association of America))
0.1
Livecodebench(UC Berkeley + MIT + Cornell (2024))
0.1
Scicode(UIUC + Argonne National Lab (2024))
0.1
Lcr(Artificial Analysis)
0.1
Aime(MAA (Mathematical Association of America))
0.1
Hle(Center for AI Safety + Scale AI (2025))
0.1
Tau2(Sierra + U Toronto + Vector Institute (2025))
0.0
Terminalbench Hard(Stanford × Laude Institute (2026))
0.0
Tau Banking
0.0
Terminalbench V2 1
0.0

LLM Stats Category Scores

(LLM Stats (zeroeval))
Instruction Following
90
Structured Output
90
Image To Text
70
Grounding
70
Math
60
Multimodal
60
Vision
60
Reasoning
50
General
50
Healthcare
50
Language
40
Legal
40
Factuality
40
Finance
40
Code
40
Physics
30
Biology
30
Chemistry
30

Pricing

Input PriceFree
Output PriceFree
Blended Price (3:1)Free

Speed

Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s

Provider Price Ranking

Provider Price Ranking

2 providers

Cheapest: OpenRouterMost Expensive: Kilo Gateway
ProviderInputOutput
1OpenRouterCheapest
$0.05
$0.1
2Kilo Gateway
$0.05
$0.1

Compare pricing across different API providers for this model.

External Sources